Personnel archive information extraction intelligent auditing method and system based on large language model

Through multimodal analysis and spatial-temporal consistency map construction based on large language model, the timestamp conflicts and attribute information mismatch of multi-source heterogeneous data in personnel file management are solved, efficient intelligent auditing and adaptive data repair are achieved, and the accuracy and audit efficiency of data consistency analysis are improved.

CN120297908APending Publication Date: 2025-07-11SHANDONG GAODI DATA SERVICE CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510413518.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing technology has data consistency problems in personnel file management, especially the time stamp conflicts of multi-source heterogeneous data, mismatch of attribute information and difficult to identify and process cross-data source conflicts, resulting in manual review taking time and inefficient and inaccurate review, which cannot meet the needs of efficient and intelligent review.

Method used

The multimodal parser based on the large language model is used to extract information from unstructured personnel files, and the timeline sequence and attribute feature set are generated through spatial and temporal features decoupling. The semantic understanding ability of the large language model is used for cross-modal alignment verification, a spatio-temporal consistency graph is built and contradiction nodes are identified, the risk conduction probability matrix and three-dimensional reliability vector are calculated, the contradiction prevention solution tree is built, and the interpretation review report is generated.

Benefits of technology

It improves the accuracy and adaptability of data consistency analysis, reduces the cost of manual audits, realizes efficient and intelligent audits, provides the traceability path and impact weight of data conflicts, and improves the conflict judgment efficiency and adaptability in complex data environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297908A_ABST
    Figure CN120297908A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information auditing management, in particular to a personnel archive information extraction intelligent auditing method and system based on a large language model, and the method comprises the following steps: generating a timeline sequence and an attribute feature set which are separated through spatial-temporal feature decoupling; performing cross-modal alignment verification on the timeline sequence and the attribute feature set, and outputting a time-space consistency map with conflict marks; generating a risk conduction probability matrix by analyzing abnormal nodes; calculating timeline credibility, attribute credibility and association credibility to form a three-dimensional credibility vector; and automatically executing logic self-consistent repair or generating an audit report comprising a contradictory traceability path. According to the method, the traceability path and the influence weight of the data conflict are provided, a visual basis is provided for manual auditing, and the conflict judgment efficiency in a complex data environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information review management, and particularly to an intelligent review method and system for personnel file information extraction based on a large language model. Background Art

[0002] In scenarios such as personnel file management, enterprise compliance review, and government administrative data processing, there are a large number of multi-source heterogeneous data from electronic documents, scanned materials, and database records. Due to different data sources, inconsistent formats, and time lags in information updates, data consistency problems are widespread, such as timestamp conflicts, mismatches in attribute information such as positions or educational backgrounds, and contradictory information across data sources that is difficult to effectively identify and process. The manual review method relies on empirical judgment, is time-consuming and inefficient, and it is difficult to form a systematic contradiction tracing and resolution strategy, unable to meet the high-efficiency intelligent review requirements in a large-scale data environment.

[0003] Existing data consistency analysis methods mainly rely on rule matching and simple feature comparison, and are difficult to handle the deep correlation relationships of cross-modal data. In the time dimension, there is a lack of an automatic detection and repair mechanism for breakpoints in time series, resulting in insufficient verification of data timeliness. At the attribute level, traditional methods are difficult to accurately measure the rationality of attribute changes and cannot identify abnormal patterns of attribute mutations. Summary of the Invention

[0004] The present invention provides an intelligent review method and system for personnel file information extraction based on a large language model, which integrates multi-modal data, constructs an efficient data consistency analysis system, has spatio-temporal consistency detection, risk conduction analysis, credibility calculation, contradiction resolution, and dynamic optimization, realizes the deep integration and intelligent review of time, attributes, and cross-source data, improves the accuracy and adaptive ability of data consistency analysis, reduces the manual review cost, and provides an interpretable review report.

[0005] The intelligent review method for personnel file information extraction based on a large language model includes the following steps: S1: An information extractor based on a large language model is used to extract information from unstructured personnel files. Based on the extracted information, a separated timeline sequence and an attribute feature set are generated through spatio-temporal feature decoupling, where the timeline sequence is a continuous temporal encoding, and the attribute feature set includes discrete vector representations of position levels and educational backgrounds; S2: The timeline sequence and the attribute feature set are subjected to cross-modal alignment verification, and the semantic understanding ability of the large language model is used to compare the logical relationships among electronic documents, scanned materials, and database records, and a spatio-temporal consistency map with conflict marks is output, and cross-document contradiction nodes are identified; S3: Construct a risk conduction chain based on the spatio-temporal consistency map, generate a risk conduction probability matrix by analyzing abnormal nodes, where the abnormal nodes include breakpoints in the timeline sequence, attribute mutation regions, and cross-source contradiction clusters; S4: Use the credibility dynamic assignment algorithm to calculate the timeline credibility (evaluating time continuity), attribute credibility (evaluating feature stability), and association credibility (evaluating multi-source consistency) respectively, and form a three-dimensional credibility vector; S5: Input the risk conduction probability matrix and the three-dimensional credibility vector into the contradiction resolution decision tree, and automatically perform logic self-consistency repair or generate an audit report including the contradiction tracing path.

[0006] Optionally, the specific content of S1 includes: S11, construction of a multi-modal parser: The multi-modal parser includes a visual branch and a text branch to achieve in-depth parsing of unstructured personnel files and fuse multi-modal information from images and texts, where: The visual branch includes a convolutional attention module: The visual branch extracts features from the images of scanned documents to obtain the spatial structure and layout information of the document content. The scanned document images go through multiple levels of convolutional operations to extract local features at different scales. The convolutional layer calculates the features of different regions through a sliding window to capture the local structural information of characters, tables, and seals. The spatial attention mechanism enhances the attention to local regions. By analyzing the importance of different region features, it calculates the attention weights and adjusts the feature representation according to the attention weights to improve the extraction accuracy of key information. For example, information such as text regions, seals, and signatures is usually more important than blank regions, so the spatial attention mechanism will assign higher weights to these regions. After convolutional and spatial attention calculations, a set of high-dimensional feature representations are output; The text branch includes a pre-trained language model based on Transformer. The text branch parses the text information in the electronic document, splits the input text into a sequence of words, and converts it into vector form through a word embedding model. At the same time, a position encoding mechanism is added to ensure that the sequential information of the text is not lost. The text data enters the pre-trained language model based on Transformer. The pre-trained language model based on Transformer uses a multi-layer self-attention mechanism to make each word pay attention to other words in the whole text and understand the context relationship. For personnel files, this mechanism is particularly important because the text content usually contains hierarchical structures and time sequences, such as a person's employment experience and educational background at different time periods. After multiple rounds of calculations, a high-dimensional feature vector is finally generated; After the visual and text features are extracted respectively, they are fused into the same feature space using the multi-head attention mechanism to establish a complete file information representation; S12, Spatiotemporal Feature Decoupling: Decompose the fused features, and extract temporal features and attribute feature sets respectively. The attribute feature set includes a position level vector and an educational background vector; S13, Feature Structured Output: Use a bidirectional long short-term memory network to model the time series data of the timeline, taking into account both past and future information. The bidirectional long short-term memory network includes a forward LSTM and a backward LSTM. The forward LSTM processes the time series from the past to the present, and the backward LSTM processes the time series from the future to the past. Integrate the information of the forward LSTM and the backward LSTM to obtain a complete timeline sequence, which includes the employment stage and the certificate validity period, and calculate the time break probability based on the complete timeline sequence.

[0007] Optionally, the specific steps of S2 are as follows: S21, Multi-source Feature Representation: Based on the timeline sequence and the attribute feature set, extract features from the electronic document source, the scanned material source, and the database record source, perform feature encoding on the data from different data sources, conduct cross-modal contrast analysis, and unify the features from all sources to the same dimension; S22, Cross-modal Semantic Alignment: Align the semantics of the data from different sources to ensure that text information and visual information can be compared in the same semantic space, including cross-modal triple matching: Calculate the semantic similarity between the electronic document, the scanned material, and the database record, construct triple relationships, quantify the information consistency between different data sources, use a cross-modal similarity calculation method to measure the matching degree between the electronic document features and the scanned material features, between the electronic document features and the database record features, and between the scanned material and the database record respectively, calculate the comprehensive similarity of the triple, calculate the difference between the feature vectors of different data sources, evaluate the matching degree, and generate a cross-modal consistency score; S23, Conflict Detection and Atlas Construction: Based on the cross-modal consistency score, detect conflicts between data sources, construct a spatio-temporal consistency atlas, mark potential data conflicts, generate a data consistency atlas, construct three types of nodes: time, attribute, and data source, and establish connections according to the data consistency relationship. Among them, the nodes with data consistency are connected by consistency edges, and the nodes with data conflicts are connected by contradiction edges, and the conflict information is recorded; S24, Contradictory Node Identification: Structurally describe the detected conflict data and generate a contradictory information mark. For each conflict node, record the conflict type, the involved data source, and the conflict confidence level, generate a contradictory description tuple, calculate the probability of the conflict occurring, and determine the conflict confidence level by combining factors such as time difference, attribute similarity, and consistency score.

[0008] Optionally, the conflict detection in S23 specifically includes: Time conflict detection: Compare the timestamps of electronic documents, database records, and the formation time points of scanned materials. If the time difference between data sources exceeds a preset threshold, it is determined as a time conflict; Attribute conflict detection: Calculate the semantic similarity of job positions, educational backgrounds, and salaries. If the similarity of corresponding fields in different data sources is lower than the preset threshold, it is determined as an attribute conflict; Cross-source contradiction analysis: Analyze the overall matching degree of information between data sources. If the overall consistency score exceeds the set threshold, it is determined as a cross-source contradiction, and the conflict details are recorded.

[0009] Optionally, the S3 specifically includes: S31, Abnormal node extraction and classification: Based on the data structure of the spatio-temporal consistency graph, extract abnormal nodes in time series, attribute features, and multi-source information matching, and classify them to support risk analysis and conduction modeling, including identifying break points in the time line sequence, detecting attribute mutation regions, and identifying cross-source contradiction clusters; S32, Risk conduction chain modeling: Based on the logical association relationships between abnormal nodes, construct risk conduction paths and establish a risk conduction topological structure; if there is an attribute mutation at the downstream time point of the time line break point, establish a conduction path from time break to attribute mutation; if the associated data sources of the cross-source contradiction cluster include attribute mutation points, establish a conduction path from cross-source contradiction to attribute mutation; based on the identified time line break points, attribute mutation regions, and cross-source contradiction clusters, and their conduction paths, construct the topological structure of the risk conduction chain to ensure that the association relationships between abnormal nodes can be reflected in the graph structure.

[0010] S33, Conduction probability calculation: Combine the historical risk database to calculate the conduction probability between abnormal nodes, and model the risk impact degree based on statistical weights; calculate the conduction weight from time break to attribute mutation based on the time interval of the time line break point and the number of attribute mutations related to this time point in historical data; calculate the conduction weight from cross-source contradiction to attribute mutation based on the number of contradiction edges of the cross-source contradiction cluster and the data source alignment score; calculate the risk conduction probability based on the conduction weight, the co-occurrence frequency of nodes in historical data, and the bias parameter, and normalize it to generate a risk conduction probability matrix for quantitative analysis of the risk propagation probability between abnormal nodes.

[0011] Optionally, the abnormal node extraction and classification in S31 specifically includes: Identifying break points in the time line sequence: Calculate the time interval between adjacent data points in the time series, determine whether the interval exceeds the time break determination threshold. If the condition is met, mark this time point as a break point and classify it into the time line abnormal node set; Attribute Mutation Region Detection: Calculate the change amplitude of the attribute feature vectors at each time point in the time series. If the change value exceeds the attribute mutation threshold, it is considered that the attribute at this time point has mutated and is classified into the set of attribute abnormal nodes; Cross-source Contradiction Cluster Identification: Count the number of contradictory edges of each node in the spatio-temporal consistency graph. If the number of contradictory edges exceeds the set number threshold, it is determined that the node is in a high-conflict state and is classified into the set of cross-source contradictory nodes.

[0012] Optionally, the S4 specifically includes: S41, Timeline Credibility Calculation: Based on the timestamp information of the timeline sequence, calculate the overall credibility of the timeline; including counting the time intervals between adjacent time periods in the timeline sequence, identifying time breakpoints, and calculating the proportion of them in the entire time span; counting the number of overlapping relationships between time periods to measure the conflict degree of time information; combining the proportion of time intervals and the number of overlaps, and calculating the timeline credibility through the set time credibility weight coefficient, and normalizing the numerical range to the credibility index; S42, Attribute Credibility Calculation: Based on the change trend of attribute features in the time series, calculate the stability of attribute data; including extracting the corresponding attribute feature vectors according to each time point in the time series, calculating the attribute change amplitude between consecutive time points, and counting the mean square deviation of the overall attribute change; according to the stability of attribute changes, calculate the attribute credibility through an exponential decay function. If a certain change exceeds the set mutation threshold, adjust the credibility evaluation; S43, Association Credibility Calculation: Based on the spatio-temporal consistency graph, calculate the association consistency credibility between data sources; including counting the number of contradictory edges in the spatio-temporal consistency graph and calculating its proportion relative to the entire data network; calculating the average value of the data source alignment scores to measure the consistency degree between multi-source data; combining the number of contradictory edges and the data alignment scores, calculating the association credibility, and adjusting the distribution of credibility values through a non-linear mapping function to make it reflect the credibility level of data consistency; S44, Three-dimensional Credibility Vector Generation: Based on the timeline credibility, attribute credibility, and association credibility, generate a three-dimensional credibility vector; including using a non-linear normalization method to convert the credibility values of the three dimensions to the same scale range.

[0013] Optionally, the S5 includes the construction of a contradiction resolution decision tree. Based on risk conduction analysis and credibility evaluation, construct a contradiction resolution decision tree to automatically handle data conflicts and generate interpretable audit results, specifically including: S51, Input the risk conduction probability matrix and the three-dimensional credibility vector; S52, Logic Self-consistent Repair: Set the logic self-consistent repair constraint conditions, and automatically correct the data conflicts that meet the self-consistent repair criteria. The automatic correction includes: Calculate the degree of breakage of the time series, infer a reasonable empty window period based on historical data, and fill in the missing time points; Evaluate the trend of attribute changes, introduce a smoothing factor and adopt a smoothing method to adjust abnormal attribute values to conform to the logical change law; S53, Audit report generation: Set the constraint conditions for audit report generation, and generate an audit report for those that meet the constraint conditions for audit report generation. The report includes: Visually display the topological mapping from the spatio-temporal consistency map to the risk conduction chain; Original evidence comparison view: Display the conflicting fields of electronic documents, scanned copies, and database records.

[0014] Optionally, the logical self-consistency repair constraint conditions are set based on the timeline credibility, attribute credibility, and risk conduction probability; the constraint conditions are set based on the risk conduction probability and association credibility.

[0015] An intelligent audit system for personnel file information extraction based on a large language model, used to implement the intelligent audit method for personnel file information extraction based on a large language model as described above, including the following modules: Multimodal parsing module: Used to extract time information and attribute features from unstructured personnel files, and generate a timeline sequence and an attribute feature set after spatio-temporal feature decoupling; Cross-modal alignment verification module: Used to compare the logical consistency of electronic documents, scanned materials, and database records, and output a spatio-temporal consistency map with conflict marks; Risk conduction analysis module: Used to construct a risk conduction chain based on the spatio-temporal consistency map and calculate the risk conduction probability matrix; Credibility evaluation module: Used to calculate the timeline credibility, attribute credibility, and association credibility, and form a three-dimensional credibility vector; Contradiction resolution and audit module: Used to construct a contradiction resolution decision tree based on the risk conduction probability matrix and the three-dimensional credibility vector, and perform logical self-consistency repair or generate an audit report including the contradiction traceability path.

[0016] Advantages of the present invention: The present invention adopts multi-source data fusion technology to uniformly model heterogeneous data such as electronic documents, scanned materials, and database records, construct a spatio-temporal consistency map to analyze time conflicts, attribute conflicts, and cross-source contradiction relationships, build a hierarchical association relationship of data anomalies through risk conduction chain modeling, quantify the conduction effects of time breaks, attribute mutations, and cross-source contradiction clusters, accurately identify the sources of data inconsistencies based on association credibility calculation and contradiction edge density analysis, and provide a traceability path and impact weight of data conflicts in combination with an interpretability report, providing an intuitive basis for manual review and improving the conflict determination efficiency in complex data environments.

[0017] The present invention adopts an adaptive correction strategy based on three-dimensional credibility calculation to fill in the time break interval, and adjusts abnormal attribute changes according to the attribute smoothing strategy to ensure that the change trends of time and attributes conform to the data evolution law. Through a repair verification mechanism, the credibility change of the corrected data is evaluated in real time, and the correction strategy is optimized by combining dynamic threshold adjustment to enhance the adaptive ability in different application scenarios and achieve high-credibility verification after data repair.

[0018] The present invention automatically adjusts the decision threshold based on the statistical analysis of long-term review data, enabling the system to dynamically optimize the time conflict determination threshold, attribute credibility evaluation criteria, and cross-modal alignment requirements according to the data scale and industry characteristics, enhancing the review adaptability in large-scale data environments, and achieving an efficient and stable intelligent data consistency review closed-loop. Brief Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only for the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 Schematic diagram of the intelligent review method process for the embodiments of the present invention; Figure 2 Schematic diagram of the system function modules for the embodiments of the present invention. Detailed Embodiments

[0021] The following will describe the present invention in detail in combination with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; and the drawings are only for more specifically describing the embodiments and are not intended to specifically limit the present invention.

[0022] It should be noted that in the specification, references to "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment includes such specific features, structures, or characteristics. Additionally, when describing a specific feature, structure, or characteristic in connection with an embodiment, implementing such feature, structure, or characteristic in connection with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0023] Generally, terms can be understood at least in part from their use in context. For example, at least in part depending on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather can alternatively, at least in part depending on the context, allow for the existence of other factors that may not be explicitly described.

[0024] As Figure 1 shown, the intelligent auditing method for personnel file information extraction based on a large language model includes the following steps: S1: Use a multi-modal parser based on a large language model to extract information from unstructured personnel files. Based on the extracted information, generate a separated timeline sequence and an attribute feature set through spatio-temporal feature decoupling. Among them, the timeline sequence is a continuous temporal encoding, and the attribute feature set includes discrete vector representations of position levels and educational backgrounds. S2: Perform cross-modal alignment verification on the timeline sequence and the attribute feature set. Utilize the semantic understanding ability of the large language model to compare the logical associations among electronic documents, scanned materials, and database records, and output a spatio-temporal consistency map with conflict marks, and identify cross-document contradiction nodes. S3: Construct a risk conduction chain based on the spatio-temporal consistency map, and generate a risk conduction probability matrix by analyzing abnormal nodes. Abnormal nodes include timeline sequence breakpoints, attribute mutation regions, and cross-source contradiction clusters. S4: Use a credibility dynamic assignment algorithm to calculate the timeline credibility (evaluating time continuity), attribute credibility (evaluating feature stability), and association credibility (evaluating multi-source consistency) respectively, and form a three-dimensional credibility vector. S5: Input the risk conduction probability matrix and the three-dimensional credibility vector into a contradiction resolution decision tree, and automatically perform logical self-consistency repair or generate an audit report including a contradiction traceability path.

[0025] S1 specifically includes: S11, construction of the multi-modal parser: Adopt a large language model with a visual-text dual-stream architecture, where: The visual branch extracts the image features of the scanned document through a convolutional attention module, and its output feature map is expressed as: , where represents the input scanned document image, is a composite operation composed of a convolutional layer and a spatial attention module; The text branch uses a pre-trained language model to parse the content of the electronic document and generate a text feature vector : , where is the input text sequence, LM is a pre-trained language model based on the Transformer architecture, represents the feature dimension (default 768 dimensions); Design a cross-modal feature fusion layer to align the image and text features through a multi-head attention mechanism to obtain the fused feature : ; S12, spatio-temporal feature decoupling: Perform a decoupling operation on the fused feature : S121, temporal feature extraction: Use a timestamp alignment algorithm to parse the time information and generate a continuous temporal encoding: , where represents the normalized timestamp (converted to Unix timestamp and normalized to the interval), represents the time span, represents the overlap ratio with the previous stage; S122, attribute feature extraction: Generate a discretized vector through hierarchical encoding: Position level vector: , where represents the rank one-hot encoding ( is the total number of preset ranks), represents the department name embedding vector (from the pre-trained matrix is the total number of departments), represents the cosine similarity calculation, is the reporting level depth (integer encoding with the root node as 0); Educational background vector: , where represents the normalized value of the college ranking, represents the professional category embedding ( is the total number of preset professional types), represents the degree level embedding is the number of degree levels); S13, feature structured output: S131, Timeline sequence generation: Use bidirectional LSTM to perform temporal modeling on the temporal encoding sequence: Forward LSTM , depending on historical information ; Backward LSTM , depending on future information ; Time break probability calculation: , where represents a trainable weight matrix, is a trainable bias term, and [;] represents the vector concatenation operation; S132, Attribute feature set construction: Map to a low-dimensional space through an embedding layer: ; where represents a trainable projection matrix, is the target embedding dimension (default 128 dimensions).

[0026] S2 specifically includes: S21, Multi-source feature mapping: Electronic document features : Extract normalized timestamps from the timeline sequence generated by S1 , and obtain the job level vector and educational background vector from the attribute feature set to construct an electronic document feature group: , represents the set of timestamps extracted from the electronic document; Scanned material features : Based on the visual branch parsing result of S1, extract the timestamps and the handwritten job attributes in the scanned document to generate a scanned feature group: , represents the set of timestamps parsed from the scanned material; Database record features : Align the database structured fields with the output of S1 and extract the timestamps and the normalized attribute vector to form a database feature group: , represents the set of timestamps extracted from the database records; S22, Calculate the triple alignment weight: , where, Represents a specific triple being currently calculated, with a value range of , corresponding to the feature vectors of the corresponding data sources , Indicates traversing all possible triples, Represents the alignment weights between three data sources. The normalization denominator ensures that the sum of all attention scores is 1. Is the cross-modal similarity: , where the dot product operation is used to measure the alignment degree between three data sources, and average normalization avoids the influence caused by different feature dimensions.

[0027] Calculate the cross-modal consistency score: , where, Represents the summation index. The weighted L2 norm calculates the distance between data sources, and finally obtains the cross-modal consistency score for subsequent conflict determination.

[0028] S23, Conflict Detection and Atlas Construction: After calculating the cross-modal consistency, judge data conflicts according to the following rules and construct a spatio-temporal consistency atlas , Represents the node set, Represents the connection edges between nodes: (1) Conflict Judgment Rules: Time Conflict: If the timestamp in the electronic document and the timestamp in the database record differ by more than the set threshold, that is: , where, Is the time threshold, default set to 7 days, considering the time tolerance of the personnel process; Attribute Conflict: If the job information in the scanned material and the job information in the database record have a semantic similarity lower than the set threshold: , where, Is the semantic embedding representation of the job field, Is the similarity threshold, set to 0.8; Cross-document Contradiction: If the cross-modal consistency score exceeds the set threshold: , where, Is set to 1.2, indicating a large inconsistency between data sources.

[0029] (2) Construct a Spatio-Temporal Consistency Atlas: Set the node set: , where, Represents time information, Represents attribute information such as job and education level, Represent data sources (electronic documents, scanned materials, databases); Set edge set: Consistency edge: ; : ; Respectively represent nodes and node . The conflicting edges are highlighted during visualization for reviewers to check.

[0030] S24, conflicting node identification : For each conflict detection result, generate a conflicting description tuple for subsequent conflict resolution and interpretive analysis: , where the value range of type is: time conflict, attribute conflict, cross-source contradiction, represents a pair of data sources where the conflict occurs, The calculation method is: , where is the trainable weight parameter, is the Sigmoid function, ensuring that the confidence value is between (0,1), and is used to calculate the conflict confidence, represents the time difference, is the cosine similarity, measuring the matching degree of attribute information (position, education level, etc.).

[0031] S3 specifically includes: S31, abnormal node extraction and classification: Extract three types of abnormal nodes from the spatio-temporal consistency graph for constructing a risk conduction chain, including: Break points in the timeline sequence: , where is the interval between adjacent time points, days, representing the time break determination threshold, represents the set of time break points; Attribute mutation area: , where is the attribute vector at time point , is the attribute vector at time point , is the Euclidean distance of the attribute feature vector, measuring the degree of attribute change, is the attribute mutation threshold, represents the set of attribute mutation areas; Cross-source contradiction cluster: , where It is a contradictory edge. If the number of contradictory edges of a certain node is greater than or equal to 2, it is determined as a cross-source contradiction cluster. Denotes the set of cross-source contradiction clusters. Is the number of contradictory edges of the current node.

[0032] S32, Risk conduction chain modeling: Based on the abnormal nodes, construct the risk conduction path and establish the topological structure, including: Temporal conduction: If there is an attribute mutation area downstream of a breakpoint in a timeline sequence, generate a conduction edge: , Is the set of conduction edges from the time breakpoint to the attribute mutation area; Cross-source conduction: If the associated electronic document or scanned material of a certain cross-source contradiction cluster contains an attribute mutation area, generate a conduction edge: , Is the set of conduction edges from the cross-source contradiction cluster to the attribute mutation area; Construct the risk conduction chain topology: , Is the risk conduction chain topological structure, including all abnormal nodes and conduction edges, where: Abnormal nodes: Include time breakpoints, attribute mutation areas, cross-source contradiction clusters; Conduction edges: Include the conduction path from time breakage to attribute mutation, and the conduction path from cross-source contradiction to attribute mutation.

[0033] S33, Conduction probability calculation: Based on the historical risk database, statistically calculate the conduction probability between abnormal nodes and calculate the risk weights, including: Temporal breakage conduction weight : , where, Is the empirical coefficient, Represents the number of attribute mutation nodes associated with the current time breakpoint; Cross-source contradiction conduction weight : , where, Is the weight coefficient, Represents the number of contradictory edges of the current cross-source contradiction cluster, Is the mean of the cross-modal alignment scores; Risk conduction probability matrix: , where, Represents the co-occurrence frequency between nodes and in the historical data, Is the trainable bias term. The Sigmoid function is used to normalize the conduction probability so that its value range is between (0,1), Represents the node to the risk conduction probability therebetween.

[0034] In S2 (Spatio-temporal Consistency Map), three thresholds , , are used to define conflicts: (Time Conflict Threshold): Used to determine whether timestamps match. If the time difference exceeds , it is determined as a time conflict.

[0035] (Attribute Conflict Threshold): Used to determine whether the semantic similarity of information such as position and education background matches. If the similarity is lower than , it is determined as an attribute conflict.

[0036] (Cross-document Contradiction Threshold): Used to measure the overall information consistency score between different data sources. If it exceeds , it indicates that there are serious inconsistencies between the data.

[0037] In S3 (Risk Conduction Chain Modeling), the definition of abnormal situations is further refined, deriving , ; (Time Break Judgment Threshold): Inherited from in S2, but focuses on the continuity within the time series; judges the time conflicts of different data sources, focuses on the gaps on the time axis; after S2 discovers a time conflict, S3 further analyzes whether this time conflict leads to a break in the time series.

[0038] (Attribute Mutation Judgment Threshold): Inherited from in S2, but focuses on the mutation points in time; is used to compare the attribute conflicts between different data sources, is used to analyze whether the attribute changes within the same data source are drastic; after S2 discovers an attribute conflict, S3 further determines whether this attribute change shows a mutation trend in time. In addition, S3 also uses count to determine cross-source contradiction clusters, which is highly correlated with (Cross-document Contradiction Threshold) in S2: is used to identify information conflicts between individual data sources, count Used to judge the overall contradiction degree among multiple data sources. When the number of contradictory edges of a certain node exceeds a certain threshold (2 or more), it indicates that a "contradiction cluster" has been formed, which has a wider influence range than a single conflict.

[0039] Based on the risk conduction probability matrix, analyze the risk propagation mode, and support risk priority assessment and decision-making support. By traversing the topological structure of the risk conduction chain, analyze the high-risk paths, and screen the risk propagation chains that may lead to systemic problems. According to the risk conduction probability matrix, predict the influence probability between different abnormal nodes, and give early warnings for high-risk nodes and paths; set the risk threshold, screen the paths whose risk conduction probability exceeds the threshold, and mark them as high-risk paths to support priority handling.

[0040] S4 specifically includes: S41, Timeline credibility calculation: Input: Timeline sequence: ; Calculation rule: , where is the interval between adjacent time periods, is the total span of the entire timeline, is the weight coefficient of time credibility, which can be dynamically adjusted according to the data scale; S42, Attribute credibility Calculation: Input attribute feature set: ; Calculation rule: , where is the attribute feature vector at time point , is the stability decay factor, , Mutation detection: If a single change satisfies: , then is directly degraded to 0.3 times, where is the mutation threshold, , is the total number of time periods of the time series, that is, the number of discrete time points in the timeline.

[0041] S43, Association credibility: Calculation: Input: From the spatio-temporal consistency map: Set of contradictory edges and alignment score ; Calculation rule: , where is the contradictory edge density: , where is the total number of nodes in the spatio-temporal consistency graph, is the penalty factor, ; S44, three-dimensional credibility vector Generate: Normalization processing: , where: is type normalization function: , is the steepness factor of the S-type function, , .

[0042] S5 includes the construction of a conflict resolution decision tree. Based on risk conduction analysis and credibility assessment, a conflict resolution decision tree is constructed to automatically handle data conflicts and generate interpretable audit results, specifically including: S51, Receive the risk conduction probability matrix and the three-dimensional credibility vector ; Decision rule: S52, Logic self-consistency repair trigger condition: , for the conflict nodes that meet the conditions, perform one of the following automatic repair operations: Intermittent fracture filling: Deduce a reasonable empty window period , is the empirical coefficient, ; Attribute mutation smoothing: Generate a transitional correction vector , is the smoothing factor, ; S53, Interpretable report generation condition: , generate an interactive report containing the following elements: Conflict traceability path graph: Visually display the topological mapping from the spatio-temporal consistency graph to the risk conduction chain; Original evidence comparison view: Side by side display the conflict fields of electronic documents, scanned copies, and database records.

[0043] The logic self-consistency repair constraint conditions are set based on the timeline credibility, attribute credibility, and risk conduction probability; the constraint conditions are set based on the risk conduction probability and the associated credibility.

[0044] As Figure 2 shown, the intelligent audit system for personnel file information extraction based on the large language model is used to implement the above intelligent audit method, including the following modules: Multimodal Analysis Module: It is used to extract time information and attribute features from unstructured personnel files, and generate a timeline sequence and an attribute feature set after decoupling spatio-temporal features; Cross-modal Alignment and Verification Module: It is used to compare the logical consistency of electronic documents, scanned materials, and database records, and output a spatio-temporal consistency map with conflict marks; Risk Propagation Analysis Module: It is used to construct a risk propagation chain based on the spatio-temporal consistency map and calculate a risk propagation probability matrix; Credibility Evaluation Module: It is used to calculate timeline credibility, attribute credibility, and association credibility, and form a three-dimensional credibility vector; Contradiction Resolution and Review Module: It is used to construct a contradiction resolution decision tree based on the risk propagation probability matrix and the three-dimensional credibility vector, perform logically self-consistent repair, or generate a review report containing the contradiction traceability path.

[0045] This invention covers any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. For the public to have a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments of this invention. However, those skilled in the art can fully understand this invention even without these detailed descriptions. Additionally, well-known methods, processes, procedures, components, and circuits, etc., are not described in detail to avoid unnecessary confusion to the essence of this invention.

[0046] The above description is only a preferred embodiment of this invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements can be made without departing from the principle of this invention, and these improvements and refinements should also be regarded as the protection scope of this invention.

Claims

1. An intelligent audit method for extracting personnel file information based on a large language model, characterized in that It includes the following steps: S1: The multimodal parser based on the large language model extracts information from the unstructured personnel files. Based on the extracted information, the separated timeline sequence and the attribute feature set are generated through spatio-temporal feature decoupling. Among them, the timeline sequence is a continuous tense encoding, and the attribute feature set includes the discrete vector representations of the position level and educational background; S2: Perform cross-modal alignment verification on the timeline sequence and the attribute feature set, and use the semantic understanding ability of the large language model to compare the logical associations among electronic documents, scanned materials, and database records, output a spatio-temporal consistency map with conflict markers, and identify cross-document contradiction nodes; S3: Construct a risk conduction chain based on the spatio-temporal consistency map, and generate a risk conduction probability matrix by analyzing abnormal nodes. The abnormal nodes include timeline sequence breakpoints, attribute mutation regions, and cross-source contradiction clusters; S4: Use the credibility dynamic assignment algorithm to calculate the timeline credibility, attribute credibility, and association credibility respectively to form a three-dimensional credibility vector; S5: Input the risk conduction probability matrix and the three-dimensional credibility vector into the contradiction resolution decision tree to automatically perform logically self-consistent repair or generate an audit report including the contradiction traceability path.

2. The intelligent audit method for extracting personnel file information based on a large language model according to claim 1, wherein The specific content of S1 includes: S11, Construction of the multimodal parser: The multimodal parser includes a visual branch and a text branch to achieve in-depth parsing of unstructured personnel files and fuse multimodal information from images and texts. Among them: The visual branch includes a convolutional attention module: The visual branch extracts features from the images of scanned documents to obtain the spatial structure and layout information of the document content. The scanned document images go through multiple levels of convolutional operations to extract local features at different scales. The convolutional layer calculates the features of different regions through a sliding window to capture the local structure information of characters, tables, and seals. The spatial attention mechanism enhances the attention to local regions. By analyzing the importance of the features of different regions, the attention weights are calculated, and the feature representation is adjusted according to the attention weights. After convolutional and spatial attention calculations, a set of high-dimensional feature representations are output; The text branch includes a pre-trained language model based on Transformer. The text branch parses the text information in the electronic document, splits the input text into a sequence of words, and converts it into a vector form through a word embedding model. At the same time, a position encoding mechanism is added. The text data enters the pre-trained language model based on Transformer. The pre-trained language model based on Transformer uses a multi-layer self-attention mechanism to make each word focus on other words in the whole text and understand the context relationship. After multiple rounds of calculations, a high-dimensional feature vector is finally generated; After the visual and text features are extracted respectively, they are fused into the same feature space by using the multi-head attention mechanism to establish a complete file information representation; S12, Spatio-temporal feature decoupling: Decompose the fused features, and extract the time features and the attribute feature set respectively. The attribute feature set includes the position level vector and the educational background vector; S13, Feature Structured Output: Use a bidirectional long short-term memory network to model the timeline data sequence to consider both past and future information simultaneously. The bidirectional long short-term memory network includes a forward LSTM and a backward LSTM. The forward LSTM processes the time series from the past to the present, and the backward LSTM processes the time series from the future to the past. The information of the forward LSTM and the backward LSTM is fused to obtain a complete timeline sequence. The timeline sequence includes the employment stage and the certificate validity period, and the time break probability is calculated based on the complete timeline sequence.

3. The intelligent audit method for extracting personnel file information based on a large language model according to claim 1, wherein The specific steps of S2 are as follows: S21, Multi-source Feature Representation: Based on the timeline sequence and the attribute feature set, extract features from the electronic document source, the scanned material source, and the database record source, perform feature encoding on the data from different data sources, conduct cross-modal comparison analysis, and unify the features from all sources to the same dimension; S22, Cross-modal Semantic Alignment: Align the semantics of the data from different sources to ensure that the text information and the visual information are compared in the same semantic space, including cross-modal triple matching: Calculate the semantic similarity between the electronic document, the scanned material, and the database record, construct a triple relationship, quantify the information consistency between different data sources, use a cross-modal similarity calculation method to measure the matching degree between the electronic document features and the scanned material features, between the electronic document features and the database record features, and between the scanned material and the database record respectively, calculate the comprehensive similarity of the triple, calculate the difference between the feature vectors of different data sources, evaluate the matching degree, and generate a cross-modal consistency score; S23, Conflict Detection and Atlas Construction: Based on the cross-modal consistency score, detect the conflicts between data sources, construct a spatio-temporal consistency atlas, mark potential data conflicts, generate a data consistency atlas, construct three types of nodes: time, attribute, and data source, and establish connections according to the data consistency relationship. Among them, the nodes with data consistency are connected by consistency edges, and the nodes with data conflicts are connected by contradiction edges, and the conflict information is recorded; S24, Contradictory Node Identification: Structurally describe the detected conflict data and generate a contradiction information mark. For each conflict node, record the conflict type, the involved data source, and the conflict confidence level, generate a contradiction description tuple, calculate the probability of the conflict occurrence, and determine the conflict confidence level by combining factors such as time difference, attribute similarity, and consistency score.

4. The intelligent audit method for extracting personnel file information based on a large language model according to claim 3, characterized in that, The conflict detection in S23 specifically includes: Time Conflict Detection: Compare the electronic document timestamp, the database record timestamp, and the scanned material formation time point. If the time difference between the data sources exceeds the preset threshold, it is determined as a time conflict; Attribute Conflict Detection: Calculate the semantic similarity of the position, education background, and salary. If the similarity of the corresponding fields in different data sources is lower than the preset threshold, it is determined as an attribute conflict; Cross-source Contradiction Analysis: Analyze the overall matching degree of the information between data sources. If the overall consistency score exceeds the set threshold, it is determined as a cross-source contradiction, and the conflict details are recorded.

5. The intelligent audit method for extracting personnel file information based on a large language model according to claim 1, wherein, The specific steps of S3 are as follows: S31, Abnormal Node Extraction and Classification: Based on the data structure of the spatio-temporal consistency graph, extract abnormal nodes in time series, attribute features, and multi-source information matching, and classify them to support risk analysis and conduction modeling, including identification of breakpoints in the timeline series, detection of attribute mutation regions, and identification of cross-source conflict clusters; S32, Risk Conduction Chain Modeling: Based on the logical association relationships between abnormal nodes, construct risk conduction paths and establish a risk conduction topology; If there is an attribute mutation at the downstream time point of the timeline breakpoint, establish a conduction path from the time break to the attribute mutation; If the associated data sources of the cross-source conflict cluster include attribute mutation points, establish a conduction path from the cross-source conflict to the attribute mutation; Based on the identified timeline breakpoints, attribute mutation regions, and cross-source conflict clusters, as well as their conduction paths, construct the topological structure of the risk conduction chain; S33, Conduction Probability Calculation: Combine with the historical risk database to calculate the conduction probability between abnormal nodes, and model the risk impact degree based on statistical weights; Calculate the conduction weight from the time break to the attribute mutation based on the time interval of the timeline breakpoint and the number of attribute mutations related to this time point in the historical data; Calculate the conduction weight from the cross-source conflict to the attribute mutation based on the number of conflict edges of the cross-source conflict cluster and the data source alignment score; Calculate the risk conduction probability based on the conduction weight, the node co-occurrence frequency in the historical data, and the bias parameter, and normalize it to generate a risk conduction probability matrix for quantitatively analyzing the risk propagation probability between abnormal nodes.

6. The intelligent audit method for extracting personnel file information based on a large language model according to claim 5, characterized in that, The abnormal node extraction and classification in S31 specifically include: Identification of Breakpoints in the Timeline Series: Calculate the time interval between adjacent data points in the time series, determine whether the interval exceeds the time break determination threshold, if the condition is met, mark this time point as a breakpoint and classify it into the timeline abnormal node set; Detection of Attribute Mutation Regions: Calculate the change amplitude of the attribute feature vectors of each time point in the time series, if the change value exceeds the attribute mutation threshold, it is considered that the attribute of this time point has mutated and classify it into the attribute abnormal node set; Identification of Cross-Source Conflict Clusters: Count the number of conflict edges of each node in the spatio-temporal consistency graph, if the number of conflict edges exceeds the set number threshold, determine that this node is in a high-conflict state and classify it into the cross-source conflict node set.

7. The intelligent audit method for extracting personnel file information based on a large language model according to claim 1, characterized in that S4 specifically includes: S41, Timeline Credibility Calculation: Based on the timestamp information of the timeline series, calculate the overall credibility of the timeline; Include counting the time intervals between adjacent time periods in the timeline series, identifying timeline breakpoints, and calculating their proportion in the entire time span; Count the number of overlapping relationships between time periods to measure the conflict degree of time information; Combine the proportion of time intervals and the number of overlaps, and calculate the timeline credibility through the set time credibility weight coefficient, and normalize the numerical range to the credibility index; S42, Attribute credibility calculation: Based on the changing trend of attribute features in the time series, calculate the stability of attribute data; including extracting the corresponding attribute feature vectors according to each time point in the time series, calculating the amplitude of attribute changes between consecutive time points, and statistically calculating the mean square error of the overall attribute changes; according to the stability of attribute changes, calculate the attribute credibility through an exponential decay function, and if a certain change exceeds the set mutation threshold, adjust the credibility evaluation; S43, Association credibility calculation: Based on the spatio-temporal consistency graph, calculate the association consistency credibility between data sources; including counting the number of conflicting edges in the spatio-temporal consistency graph and calculating its proportion relative to the entire data network; calculating the average value of the data source alignment scores to measure the consistency degree between multi-source data; combining the number of conflicting edges and the data alignment scores to calculate the association credibility, and adjusting the distribution of credibility values through a non-linear mapping function to make it reflect the credibility level of data consistency; S44, Generation of three-dimensional credibility vector: Based on the timeline credibility, attribute credibility, and association credibility, generate a three-dimensional credibility vector; including using a non-linear normalization method to convert the credibility values of the three dimensions to the same scale range.

8. The intelligent audit method for extracting personnel file information based on a large language model according to claim 1, characterized in that The above S5 includes the construction of a conflict resolution decision tree. Based on risk conduction analysis and credibility evaluation, construct a conflict resolution decision tree to automatically process data conflicts and generate an interpretable audit result, specifically including: S51, Input the risk conduction probability matrix and the three-dimensional credibility vector; S52, Logic self-consistency repair: Set the logic self-consistency repair constraint conditions, and automatically correct the data conflicts that meet the self-consistency repair criteria. The automatic correction includes: Calculating the break degree of the time series, speculating the reasonable empty window period based on historical data, and filling in the missing time points; Evaluating the trend of attribute changes, introducing a smoothing factor and using a smoothing processing method to adjust abnormal attribute values to make them conform to the logical change law; S53, Audit report generation: Set the audit report generation constraint conditions, and generate an audit report for those that meet the audit report generation constraint conditions. The report includes: Visual display of the topological mapping from the spatio-temporal consistency graph to the risk conduction chain; Original evidence comparison view: Display the conflicting fields of electronic documents, scanned copies, and database records.

9. The intelligent audit method for extracting personnel file information based on a large language model according to claim 8, characterized in that, The above logic self-consistency repair constraint conditions are set based on the timeline credibility, attribute credibility, and risk conduction probability; the constraint conditions are set based on the risk conduction probability and the association credibility.

10. An intelligent audit system for extracting personnel file information based on a large language model, used to implement the intelligent audit method for extracting personnel file information based on a large language model described in any one of claims 1-9, characterized in that, It includes the following modules: Multi-modal parsing module: Used to extract the time information and attribute features from the unstructured personnel files, and generate a timeline sequence and an attribute feature set after spatio-temporal feature decoupling; Cross-modal alignment verification module: Used to compare the logical consistency of electronic documents, scanned materials, and database records, and output a spatio-temporal consistency graph with conflict marks; Risk conduction analysis module: Used to construct a risk conduction chain based on the spatio-temporal consistency graph and calculate the risk conduction probability matrix; Credibility evaluation module: Used to calculate the timeline credibility, attribute credibility, and association credibility, and form a three-dimensional credibility vector; Contradiction Resolution and Audit Module: Used to construct a contradiction resolution decision tree based on the risk conduction probability matrix and the three-dimensional credibility vector, and perform logically self-consistent repair or generate an audit report containing the contradiction traceability path.

Citation Information

Cited By

  • Talent background investigation method based on multi-source data evaluation

    CN120598436A

  • Multi-source heterogeneous data collection and fusion method and system for environmental governance industry

    CN120873998A

  • Full-period electronic management system for immigrant archives

    CN120892412A

  • Intelligent news content error correction system based on large model

    CN120893427A

  • Dynamic credibility quantitative modeling method for cross-domain heterogeneous data

    CN121009076A