File information management method based on full life cycle

By establishing a data index library and using link prediction algorithms, the correlation between case file information is identified, solving the problem of unified management of case file information throughout its entire lifecycle and achieving efficient correlation analysis and management.

CN122019634APending Publication Date: 2026-05-12中国共产党山东省委员会党校
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
中国共产党山东省委员会党校
Filing Date
2026-01-28
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, case file information lacks unified management and correlation analysis throughout its entire lifecycle, making it difficult to discover and identify potential relationships between related entities in a timely manner, thus affecting management efficiency and accuracy.

Method used

Collect case file information throughout the entire lifecycle, establish a data index library, identify relationships and construct an initial topology, use link prediction algorithms to analyze weakly related areas, identify nodes at risk of breakage, update the topology through implicit relationship mining, and form a lifecycle-wide related topology.

Benefits of technology

It has enabled unified management and correlation analysis of case file information, reduced retrieval omissions, enhanced traceability, review and maintenance capabilities, supported screening of key targets, and improved management efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019634A_ABST
    Figure CN122019634A_ABST
Patent Text Reader

Abstract

The invention discloses a file information management method based on a full life cycle, and particularly relates to the technical field of file management. The method comprises the following steps: collecting file information generated in a full life cycle of a file, establishing a structured file information index database, identifying an initial association relationship between file association subjects, and constructing an association relationship initial topological structure; analyzing the initial topological structure of the association relationship by adopting a link prediction algorithm, identifying a weak association topological region appearing over time, judging a fracture risk node in the topological structure, and predicting an association degradation trend of a file association subject; performing association supplementary analysis on the fracture risk nodes, and mining potential implicit association between file information; and finally, updating and reconstructing the initial topological structure according to the implicit association mining data. According to the invention, accurate analysis and prediction of the association relationship of the file information can be realized, and the integrity and the intelligent level of file information management are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of case file management technology, and more specifically, to a method for case file information management based on the entire lifecycle. Background Technology

[0002] In existing technologies, case file information management is generally carried out in an independent and decentralized manner. Case file information generates a large amount of various types of information, such as text, images, and electronic data, from case filing, investigation, handling to final archiving. This information is often stored in their respective independent business systems or departments, lacking unified data association and analysis methods, making it difficult to identify and effectively mine the inherent connections between case file information in a timely manner.

[0003] Therefore, the existing technology has the following technical problem: the lack of an effective method for unified management and correlation analysis of case file information throughout its entire lifecycle makes it difficult to discover and identify potential relationships between related entities in case files in a timely manner, thus restricting the efficiency and accuracy of case file information management and analysis. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a case file information management method based on the entire life cycle to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: A method for managing case file information based on the entire lifecycle includes the following steps: S1: Collect case file information generated at different time points throughout the entire life cycle, establish a data index according to the case file information type, and generate a case file information index library; S2: Based on the case file information index, identify the initial relationships between the related entities in the case file information and establish the initial topology of the relationships; S3: Use the link prediction algorithm to analyze the initial topology of the association relationship, identify the weak association regions of the topology that appear over time, and obtain the distribution data of the weak association topology regions. S4: Based on the distribution data of weakly associated topological regions, identify the nodes at risk of breakage in the initial topological structure of the association relationship, and generate a list of predicted degradation of the association of case file related subjects. S5: Based on the node association degradation prediction list, perform association supplement analysis on the nodes at risk of breakage, identify potential implicit associations between case file information, and generate implicit association mining data; S6: Update and reconstruct the initial topology of the association based on the implicit association mining data to form the full life cycle association topology of the case file information.

[0006] In a preferred embodiment, S1 specifically refers to: Collect case file information generated at different time points throughout the entire life cycle and write it into the original case file information set. Standardize the format and unify the character encoding of the case file information and add time stamps and source stamps. Extract case file information type tags based on case file information type mapping rules; Based on the case file information type tags, time markers, and source markers, construct index keys and generate index entries; The aggregated index entries form a case file information index database.

[0007] In a preferred embodiment, S2 specifically refers to: Based on the case file information index database, index entries are read and index keys are parsed to obtain case file information type tags, time stamps, and source stamps; Extract the identifiers of the related entities from the case file information and normalize them to form a set of related entities for the case file. Generate association records based on the co-occurrence and temporal continuity of related subjects in different index entries within the same case file; Aggregate relationship records to construct the initial topology of the relationships.

[0008] In a preferred embodiment, S3 specifically refers to: Based on the initial topology structure of the association relationship, the case file association subject set and association relationship record are read, and the time window is divided according to the time mark and the time window topology structure is generated. In the time window topology, the link prediction algorithm is used to calculate the link prediction score of the case file related subject pairs, forming a link prediction score set; Based on the link prediction score set and the actual link status in the time window topology, the link deviation value is calculated to generate link deviation feature data; Based on the link deviation feature data, weakly correlated node clusters are aggregated according to the weak correlation determination rules, and weakly correlated topological region distribution data is output.

[0009] In a preferred embodiment, S4 specifically refers to: Based on the distribution data of weakly associated topological regions, weakly associated node clusters are read and mapped to the initial topological structure of the association relationship, and the case file association subject set corresponding to the weakly associated node clusters is extracted; In the initial topological structure of the association relationship, the node degree value and intermediate value of the case file association subject set are calculated to form node structure feature data; Calculate node breakage risk values ​​and generate a set of nodes at breakage risk based on node structure feature data and link deviation feature data; Based on the set of fracture risk nodes, output a list of predicted degradation of the associated entities in the case file.

[0010] In a preferred embodiment, S5 specifically refers to: Based on the case file association subject association degradation prediction list, read the set of break risk nodes and locate the index entries corresponding to the set of break risk nodes, extract the case file information associated with the index entries to form a candidate case file information set; In the candidate case file information set, perform similar merging based on case file information type tags and time sorting based on time markers to generate a candidate case file information sequence; Perform keyword matching and citation relationship parsing on the candidate case file information sequence to generate citation relationship records; Consistency checks are performed on reference records and association records to output implicit association mining data.

[0011] In a preferred embodiment, S6 specifically refers to: Based on the implicit association mining data, extract the case file association subject identifier pairs corresponding to the implicit association mining data and generate supplementary association relationship records; The supplementary relationship records are merged into the relationship records, and the initial relationship topology is updated based on the supplementary relationship records to obtain the updated topology; In updating the topology, redundant links are eliminated and connectivity checks are performed based on the case file association subject set to generate a reconstructed link set; The updated topology is reconstructed based on the reconstructed link set to obtain the full lifecycle association topology of case file information.

[0012] The technical effects and advantages of the case file information management method based on the entire life cycle of this invention are as follows: By uniformly collecting and indexing case file information throughout its entire lifecycle, a case file information index library is formed, reducing retrieval omissions caused by scattered storage. Based on the case file information index library, related entities of case files are extracted and an initial topological structure of related relationships is constructed to achieve a structured presentation of case file information relationships. The link prediction algorithm is used to output the distribution data of weakly related topological regions, and time windows are divided by time stamps for comparison to locate weak positions in the related chains. By identifying nodes at risk of breakage and forming a list of predicted degradation of related entities in case files, key object screening is supported. By tracing back the case file information associated with nodes at risk of breakage and parsing the reference relationships, implicit link mining data is generated. By incorporating the implicit link mining data and reconstructing the link set, and maintaining the consistency of the topological structure through redundant link elimination and connectivity verification, a full lifecycle related topological structure of case file information is formed, enhancing traceability, review, and continuous maintenance capabilities. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of a case file information management method based on the entire life cycle according to the present invention. Detailed Implementation

[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0015] Example Figure 1 This invention presents a method for managing case file information based on the entire lifecycle, which includes the following steps: S1: Collect case file information generated at different time points throughout the entire life cycle, establish a data index according to the case file information type, and generate a case file information index library; S2: Based on the case file information index, identify the initial relationships between the related entities in the case file information and establish the initial topology of the relationships; S3: Use the link prediction algorithm to analyze the initial topology of the association relationship, identify the weak association regions of the topology that appear over time, and obtain the distribution data of the weak association topology regions. S4: Based on the distribution data of weakly associated topological regions, identify the nodes at risk of breakage in the initial topological structure of the association relationship, and generate a list of predicted degradation of the association of case file related subjects. S5: Based on the node association degradation prediction list, perform association supplement analysis on the nodes at risk of breakage, identify potential implicit associations between case file information, and generate implicit association mining data; S6: Update and reconstruct the initial topology of the association based on the implicit association mining data to form the full life cycle association topology of the case file information.

[0016] S1: Collect case file information generated at different time points throughout the entire lifecycle, establish data indexes according to case file information types, and generate a case file information index library, including: Collect case file information generated at different time points throughout the entire life cycle and write it into the original case file information set. Standardize the format and unify the character encoding of the case file information and add time stamps and source stamps. Case file information refers to all information generated throughout the entire process of a case, from filing, investigation, and handling to archiving and reuse. This includes, but is not limited to, various types of information such as text, images, audio recordings, and electronic data, such as investigation reports, interrogation transcripts, meeting minutes, case handling decisions, audit reports, and video recordings. For case file information, automated or manual methods are used to retrieve information from the business terminals or systems of auditing, disciplinary inspection departments, or other relevant departments through pre-configured information collection interfaces or specific data upload tools. Each piece of case file information is recorded in the original case file information set without any filtering or deletion to ensure the integrity and originality of the information.

[0017] During format standardization, the original format category of the case file information is identified. For text-based case file information, it is uniformly converted to UTF-8 encoded plain text format; for image-based case file information, it is uniformly converted to specified standard video or image formats, such as MP4 or JPEG; for audio-based case file information, it is uniformly converted to standard audio formats such as MP3; and for electronic data-based case file information, it is uniformly converted to a parsable structured data format, such as JSON. For character encoding, all case file information containing text content uniformly adopts UTF-8 character encoding to avoid data reading or parsing errors caused by encoding differences. After format standardization and character encoding unification are completed, a time stamp and source stamp are added to each piece of case file information. The time stamp refers to a standardized timestamp generated based on the actual date and time of generation or recording of each piece of case file information, for example, it can be set to the format "YYYY-MM-DD HH:mm:ss", with seconds as the smallest unit of time. The source stamp refers to the name of the organization from which the case file information was generated or collected, such as an audit department or disciplinary inspection department, to facilitate traceability analysis and information tracking.

[0018] Extract case file information type tags based on case file information type mapping rules; Case file information type mapping rules refer to a predefined set of information type determination criteria used to uniformly classify case file information of different formats, contents, and uses under corresponding type labels; the type label identifies the purpose and classification of the case file information. First, the case file information, after being processed for format standardization and character encoding, is read one by one. Based on the content and format characteristics of the case file information, such as keywords in the document title, specific phrases or structural features in the document body, file extensions, information sources, and other dimensions, it is compared one by one with the predefined feature matching rules in the case file information type mapping rules. Case file information that matches successfully is assigned the corresponding case file information type label, such as investigation records, audit reports, video evidence, case conclusions, etc. If multiple matching labels appear, the label priority rule preset in the case file information type mapping rules is used to determine the final case file information type label, ensuring that each piece of case file information is assigned only a unique type label.

[0019] Based on the case file information type tags, time markers, and source markers, construct index keys and generate index entries; An index key is a composite structured string used for quickly retrieving and locating case file information. An index key consists of three parts concatenated in a specific order: a case file information type label, a time stamp, and a source stamp. These parts are separated by specific symbols, such as underscores. For example, an index key could be represented as "Investigation Record_2025-01-20 10:30:00_Audit Department". Based on this definition, each piece of case file information with an assigned case file information type label is read sequentially. The three attribute values ​​are extracted and concatenated in order to form the index key string. Each index key corresponds uniquely to one normalized piece of case file information. After the index key is constructed, the index entry generation process involves combining the constructed index key with the unique storage location identifier of the corresponding case file information, such as a file path or data storage address, into a key-value pair structure, enabling quick retrieval or location of the original case file information using the index key.

[0020] The aggregated index entries form a case file information index database; The case file information index is a structured data collection that uniformly stores all index entries. It can be configured using a relational or non-relational database. All completed index entries are sequentially written into and stored in the case file information index. Each index entry in the index uses its index key as the primary key for retrieval and the storage location of the corresponding case file information as the data content item. The storage structure uses three key attributes—case file information type tag, time stamp, and source tag—as the index's key fields to quickly complete queries and retrievals with complex combinations of conditions, achieving efficient access to case file information.

[0021] Once the case file information index is built, when it is necessary to retrieve case file information of a certain type or source, the required case file information can be quickly located and directly accessed by combining any one or more conditions among the case file information type tags, time stamps, and source tags.

[0022] S2: Based on the case file information index, identify the initial relationships between related entities in the case file information, and establish the initial topology of the relationships, including: Based on the case file information index database, index entries are read and index keys are parsed to obtain case file information type tags, time stamps, and source stamps; In the case file information index, each index entry consists of an index key and a unique storage location identifier for the corresponding case file information. The index key is a composite structured string composed of three parts: a case file information type tag, a timestamp, and a source tag, sequentially concatenated and separated by a uniform underscore. During index key parsing, the entire case file information index is traversed one by one to read all stored index entries. After reading, for each index entry, a string splitting operation is performed using the underscores in the index key as delimiters, splitting the index key into three separate attributes: the case file information type tag, the timestamp, and the source tag. The resulting case file information type tag clarifies the category of the case file information, such as audit reports, investigation records, video evidence, or case conclusions; the timestamp records the generation time of the case file information, using a uniform standard format, such as YYYY-MM-DD HH:mm:ss, accurate to the second; and the source tag indicates the institution that generated or collected the case file information, such as an auditing department or a disciplinary inspection department.

[0023] Extract the identifiers of the related entities from the case file information and normalize them to form a set of related entities for the case file. The associated subject of a case file refers to the unique identifier of the person, event, unit, or organization associated with the case file information, including but not limited to the person's name, position or number, unit's name, event number, or specific name. First, the storage location of the case file information is located based on the case file information type tag, time stamp, and source tag obtained by parsing the index key. Then, the case file information content stored in the original case file information set is read line by line. For text-based case file information, natural language processing technology is used for keyword recognition and entity extraction to extract the associated subject identifiers such as the person's name, unit's name, and event number mentioned in the case file information. For image or audio-based case file information, speech recognition technology is used to convert the audio content into text, and then natural language processing methods are used for entity extraction to obtain the associated subject identifiers. For electronic data-based case file information, structured data parsing is used to directly extract the subject identifiers that are clearly associated with the case file information. After obtaining all associated entity identifiers, a normalization process is performed, including standardizing the writing conventions of entity identifiers. For example, for names, the standard format of "surname + given name" is used; for organizations, the standard full registered name of the organization is used; and for event numbers, the standard event number format is used to avoid duplication or ambiguity caused by differences in expression for the same entity. All normalized case file associated entity identifiers are then merged and stored as a case file associated entity set.

[0024] Generate association records based on the co-occurrence and temporal continuity of related subjects in different index entries within the same case file; Co-occurrence of related entities in the same case file refers to the relationship formed when multiple related entities in the same case file information or different case file information generated at the same time point appear simultaneously. Temporal continuity refers to the temporal continuity relationship exhibited by the same related entity appearing continuously or frequently in case file information at different times. First, the set of related entities in the case file is traversed, and for each related entity identifier, other related entity identifiers with co-occurrence or temporal continuity relationships are identified in the case file information corresponding to all indexed entries. When identifying co-occurrence relationships, for each case file information, all related entity identifiers extracted from the case file information are compared one by one. When multiple different entities are found to exist simultaneously, the relationship between the entities is recorded as a co-occurrence relationship. When identifying temporal continuity relationships, based on the standardized time stamps attached to the case file information, the occurrence of the same related entity at different time points is analyzed. When the entity appears frequently at consecutive time points or within short time intervals, it is recorded as a temporal continuity relationship. The criterion for determining a temporally continuous relationship is that the time interval does not exceed a set threshold. The method for determining the set threshold is to analyze the statistical regularity of the frequency of the subject's behavior within the case file's life cycle, for example, it can be set to 7 days or 30 days. Co-occurrence relationships and temporally continuous relationships are recorded respectively, including the identifier of the associated subject, the identifier of the other associated subject, the relationship type (co-occurrence or temporally continuous), and the time mark of occurrence, forming a standardized relationship record.

[0025] Aggregate relationship records to construct the initial topology of the relationships; This system aggregates all relationship records, including all co-occurrence and temporally continuous relationships. Based on graph theory's topology approach, it constructs an initial topology structure using each case file associated entity identifier as a topology node and the relationships within the relationship records as topology edges. Each topology node in this initial structure is a unique case file associated entity identifier. The topology edges between nodes are defined by the relationship records, and each edge is labeled with its type and temporal attribute. For example, the attributes of each edge record whether it's a co-occurrence or temporally continuous relationship, and they are marked with standardized time stamps for the first and latest occurrences of the relationship, enabling temporal location and dynamic change analysis of the relationships. The initial topology structure reflects the verifiable initial relationship status among all case file associated entities.

[0026] S3: The link prediction algorithm is used to analyze the initial topology of the association relationships, identify weakly associated regions that appear over time, and obtain the distribution data of weakly associated topology regions, including: Based on the initial topology structure of the association relationship, the case file association subject set and association relationship record are read, and the time window is divided according to the time mark and the time window topology structure is generated. Each topological node in the initial topological structure of the association corresponds to a case file association entity identifier in the case file association entity set. Topological edges are defined by the association record, and each topological edge records the association type (co-occurrence relationship or temporally continuous relationship) and the time stamp of the association occurrence. Based on the standardized time stamps recorded in the association record, the method for dividing the time window is determined. The method for dividing the time window is as follows: First, determine the start and end times within the entire lifecycle. The start time is the earliest time stamp in the association record, and the end time is the latest time stamp. Select an appropriate window size based on actual management needs and data analysis granularity, such as one month or three months. According to the selected window size, the time window from the start time to the end time is continuously divided into several non-overlapping continuous time windows. Each time window is identified by its start and end times. For example, if the start time is 2024-01-01 00:00:00 and the window size is one month, then the end time of the first window is 2024-01-31 23:59:59, and so on. Within each time window, records with timestamps falling within the time window range are filtered and aggregated to form the corresponding time window topology. For each time window, the associated entity identifier is extracted to form a topology node, and the associated relationship is extracted to form a topology edge, thus obtaining the time window topology. Each time window topology only includes the associated entities and relationships that actually exist within the corresponding time period of the time window.

[0027] In the time window topology, the link prediction algorithm is used to calculate the link prediction score of the case file related subject pairs, forming a link prediction score set; Link prediction algorithms calculate the probability or strength of a potential association between any two related entities in a case file by analyzing the historical relationships between topological nodes within a time window topology. The input data for link prediction algorithms consists of data from the current time window topology and historical time window topologies. Historical time window topologies refer to one or more time window topologies preceding the current one. The algorithm predicts potential associations between topological nodes according to predefined rules. For example, using the common neighbor count algorithm, the number of neighboring nodes shared by any two topological nodes in the historical time window topology is calculated; this number represents the link prediction score between the node pairs. If the Jaccard similarity coefficient algorithm is used, the normalized link prediction score is obtained by dividing the number of common neighbors of the node pairs by the union of the number of neighbors of each node pair. If the Adamic-Adar algorithm is used, common neighbors are assigned a weight based on node degree, and the link prediction score is obtained by summing these weights. The selection method for the link prediction algorithm is as follows: First, based on historical data analysis, the sparsity of node connection patterns and the distribution characteristics of node degree are analyzed. When the sparsity is high and the differences in node degree distribution are large, the Adamic-Adar algorithm can be preferred; when the node connection patterns are relatively uniform, the Jaccard similarity coefficient algorithm or the common neighbor count algorithm can be preferred. The calculated link prediction score is a non-negative real number. The larger the link prediction score, the higher the probability that there is a correlation between the predicted node pairs. The link prediction scores of all node pairs are summarized to form a link prediction score set. Each record in the set includes the node pair identifier and the corresponding prediction score.

[0028] Based on the link prediction score set and the actual link status in the time window topology, the link deviation value is calculated to generate link deviation feature data; Link deviation refers to the degree of deviation between the strength of the association predicted by the link prediction algorithm and the actual association between node pairs in the current time window topology. The calculation method is as follows: For each node pair, firstly, the prediction score of the node pair is extracted based on the link prediction score set; then, it is confirmed whether there is an actual topological edge between the node pairs in the current time window topology. If an actual topological edge exists, the actual link state value is set to 1; otherwise, the actual link state value is set to 0.

[0029] The formula for calculating link deviation is: ; in, Represents node pairs and Link deviation value; Represents node pairs and This represents the link prediction score after normalization. Represents node pairs and The actual link state value; the normalization method is: divide the link prediction score of all node pairs in the link prediction score set by the maximum value of the prediction score in the set, so that the normalized link prediction score falls within the range of 0 to 1.

[0030] After calculating the link deviation values ​​of all node pairs, the link deviation feature data is summarized, which includes node pair identification, link prediction score, actual link status value and link deviation value.

[0031] Based on the link deviation feature data, weakly associated node clusters are aggregated according to the weak association determination rules, and weakly associated topological region distribution data is output. The weak association determination rule refers to the standard used to identify weak associations between nodes. The determination rule is expressed as a link deviation value greater than a set threshold and an actual link status value of 0, indicating a node pair predicted to have a relationship but not actually established. The link deviation threshold is determined as follows: First, the overall distribution characteristics of node pair link deviation values ​​within a historical time window are statistically analyzed. Based on the statistical data, the high percentile value is selected as the threshold for the link deviation value, for example, it can be set to the top 5% or top 10% quantile of the deviation value distribution. After identifying node pairs that meet the weak association determination rule, weak association node clusters are formed based on these node pairs. A weak association node cluster refers to the set of nodes that are predicted to have a relationship but have not actually established one. Each node cluster records the corresponding node identifier and link deviation characteristic data. The method for setting the number of nodes within a node cluster is as follows: Using node pairs that satisfy the weak association determination rules as seeds, a node clustering operation is performed. Clustering is based on the link deviation distance between nodes. For example, when using a hierarchical clustering algorithm, a clustering tree is constructed based on the link deviation distance, and then the clustering tree is pruned according to a distance threshold to obtain node clusters. The method for determining the distance threshold is the same as the method for determining the link deviation threshold. After the node clusters are aggregated, the distribution data of the weak association topology region is output, including the node identifier of each node cluster, the link deviation value between nodes, and the time window information of the node cluster, indicating the time period and the weakening state of the association relationship of the node cluster.

[0032] S4: Based on the distribution data of weakly correlated topological regions, identify the nodes at risk of breakage in the initial topological structure of the correlation relationship, and generate a list of predicted degradation of the correlation between case file subjects, including: Based on the distribution data of weakly associated topological regions, weakly associated node clusters are read and mapped to the initial topological structure of the association relationship, and the case file association subject set corresponding to the weakly associated node clusters is extracted; Weakly associated node clusters are sets of multiple case file associated entities identified based on link deviation feature data according to weak association judgment rules. Each weakly associated node cluster corresponds to a set of case file associated entities predicted by a link prediction algorithm to have a high probability of association but which have not formed an association relationship in the actual topology structure. The node identifiers of each weakly associated node cluster in the weakly associated topology distribution data are read one by one. The node identifiers clearly correspond to the case file associated entity identifiers in the case file associated entity set. All node identifiers in the read weakly associated node clusters are mapped to the initial topology structure of the association relationship, accurately locating the topology nodes in the initial topology structure that are completely consistent with the node identifiers. Based on these topology nodes, the set of case file associated entities corresponding to the weakly associated node clusters is obtained. The mapping method is as follows: using the node identifiers of the weakly associated node clusters as the retrieval basis, all topology node identifiers are compared one by one in the initial topology structure of the association relationship to determine and extract the corresponding set of case file associated entity identifiers, forming the set of case file associated entities to be analyzed.

[0033] In the initial topological structure of the association relationship, the node degree value and intermediate value of the case file association subject set are calculated to form node structure feature data; The node degree value is defined as the number of directly connected nodes of each case file-related entity's corresponding topological node in the initial topological structure of the association relationship, i.e., the number of topological edges directly connected to the node. The median value is defined as the frequency with which each topological node is on the shortest path between other node pairs within the initial topological structure of the association relationship, used to characterize the node's critical or central role in the entire topological structure. Based on the initial topological structure of the association relationship, the node degree value is calculated for each topological node corresponding to each case file-related entity in the case file-related entity set, and the total number of topological edges directly connected to the node is counted as the node degree value. Taking all topological nodes within the initial topological structure of the association relationship as the starting and ending points, all shortest paths between each node are calculated. The frequency of each case file-related entity node on the shortest path is summed and divided by the total number of shortest paths to obtain the normalized median value. The calculation of node degree value and median value can be accurately implemented using existing graph algorithms, such as using the Brandes algorithm to calculate the median value by traversing the topological nodes and calculating the shortest paths between all nodes, thus quickly obtaining the median value; the node degree value is directly obtained by counting the number of topological edges connected to the topological nodes. After the calculation is completed, the node structure feature data records the node degree value and intermediate value of the node corresponding to each case file associated subject.

[0034] Calculate node breakage risk values ​​and generate a set of nodes at breakage risk based on node structure feature data and link deviation feature data; The node breakage risk value is used to assess the degree of risk that the topological node corresponding to each case file's associated subject may experience a breakage or weakening of the association relationship in the initial topological structure of the association relationship. The calculation process for the node breakage risk value is as follows: For any topological node corresponding to a breakage risk node identifier, firstly, the node degree value and median value are read from the node structural feature data, and normalized by dividing the node degree value by the maximum value of all topological node degree values ​​and the median value by the maximum value of all topological node median values, respectively, to obtain the normalized node degree value and normalized median value; secondly, all link deviation values ​​formed by the breakage risk node identifier and other topological nodes are aggregated from the link deviation feature data, and the arithmetic mean of the link deviation values ​​is calculated to obtain the node link deviation mean; then, the normalized node degree value, the normalized median value, and the node link deviation mean are multiplied by the weight coefficient, respectively, and the three products are summed to obtain the node breakage risk value. The weight coefficient is determined by constructing a multi-factor linear regression analysis based on historical correlation records and historical link deviation feature data, and assigning weights according to the normalization of the absolute value of the regression coefficient. For example, the weight coefficients obtained by normalizing the regression coefficients can be set to 0.4, 0.4, and 0.2. The calculated node breakage risk value is a real number between 0 and 1. The higher the node breakage risk value, the higher the risk of breakage of the related subject of the case file corresponding to the node. A threshold for the node breakage risk value is set as the criterion for determining risk nodes. The threshold is determined based on the distribution characteristics of the node breakage risk value. For example, the threshold can be set as the breakage risk value of the top 10% of nodes in the node breakage risk value ranking. Nodes with a breakage risk value higher than the threshold are identified as breakage risk nodes, forming a set of breakage risk nodes.

[0035] Output a list of predicted degradation of case file related entities based on the set of fracture risk nodes; The case file association degradation prediction list records information on case file association entities with a high risk of association degradation or breakage. First, node identifiers are read one by one from the set of breakage risk nodes. Each node identifier corresponds to the name or number of the case file association entity in the set. Each breakage risk node identifier is then marked with a corresponding breakage risk value, and the nodes are sorted in descending order of their breakage risk values ​​to form the association degradation prediction list. Each entry in the association degradation prediction list includes the case file association entity identifier and the node breakage risk value, along with the node degree value, median value, and link deviation value. During the entire lifecycle management of case files, the association degradation prediction list identifies entities that may experience weakening or breakage of their associations, characterizing the stability and degradation risk of case file association entities in the association relationship topology.

[0036] S5: Based on the node association degradation prediction list, perform supplementary association analysis on nodes at risk of breakage, identify potential implicit associations between case file information, and generate implicit association mining data, including: Based on the case file association subject association degradation prediction list, read the set of break risk nodes and locate the index entries corresponding to the set of break risk nodes, extract the case file information associated with the index entries to form a candidate case file information set; Based on the case file association degradation prediction list, the set of break risk nodes is read and the corresponding index entries are located. First, the break risk node identifier in each break risk node set in the case file association degradation prediction list is read one by one. The break risk node identifier corresponds to the topological node in the initial topology of the association relationship, and also corresponds to the case file association entity identifier in the case file association entity set. After reading the break risk node identifier, the case file information associated with each index entry is compared in the case file information index database. The index entries containing the case file association entity identifier corresponding to the break risk node identifier are retrieved and located. The retrieval process includes traversing all index entries in the case file information index database, reading the case file information associated with each index entry according to the case file information type label, time stamp, and source stamp recorded in the index key, and identifying whether there is a case file association entity identifier corresponding to the break risk node identifier through defined natural language processing technology or structured data analysis methods. If the identification result is that it exists, it is determined that the location is successful and the index entry is recorded. All successfully located index entries are aggregated and the case file information associated with each index entry is extracted one by one to form a candidate case file information set. The candidate case file information set contains all case file information directly associated with the case file associated subject identifier corresponding to the set of fracture risk nodes.

[0037] In the candidate case file information set, perform similar merging based on case file information type tags and time sorting based on time markers to generate a candidate case file information sequence; The system reads the case file information type tags attached to each piece of information in the candidate case file information set, one by one. These tags are generated according to case file information type mapping rules and are used to identify the category of the case file information, including but not limited to type tags for investigation records, audit reports, video evidence, and case conclusions. After reading all the case file information type tags, based on the consistency judgment rules, case files with the same case file information type tags are grouped into a single subset of the same type. All case file information type tags within each subset of the same type are completely identical, facilitating internal retrieval. After grouping by type, for each subset of the same type of case file information, it is sorted according to the standardized timestamps attached to the case file information, based on the actual generation or recording time order of the case file information. The sorting rule is to arrange the case file information within all subsets of the same type of case file information in the candidate case file information set in ascending order of timetamps from earliest to latest, generating a candidate case file information sequence.

[0038] Perform keyword matching and citation relationship parsing on the candidate case file information sequence to generate citation relationship records; Keyword matching refers to using natural language processing methods to scan each piece of case file information within a candidate case file information sequence. By using a pre-defined set of keywords, such as names of persons, organizations, case file association identifiers, case numbers, or investigation matters, keywords reflecting relationships within the case file information are extracted. Citation relationship analysis refers to using semantic analysis and text structure recognition methods to analyze citation or indirect citation relationships between case file information within a candidate case file information sequence. A citation relationship is when one case file information contains a textual reference to the title or number of another case file information. An indirect citation relationship is when keyword matching and semantic analysis determine a content-based or factual connection between case file information, but no citation is explicitly stated. The citation relationship parsing analyzes the text content of candidate case file information sequence line by line. Based on the semantic structure of citation expression patterns such as "see...report" and "based on...meeting minutes", it extracts and records the unique identifier of the cited case file information through rule base matching. For example, the unique identifier is a combination of case file information type label and time stamp. In this way, an accurate and verifiable citation relationship record is established between the cited case file information. Each record in the citation relationship record identifies the source and target case file information of the citation relationship, the citation type (direct citation or indirect citation), and the location of the cited text.

[0039] Consistency checks are performed on reference relationship records and association relationship records to output implicit association mining data; The unique identifiers of the source and target case file information recorded in each reference relationship record are read one by one. Then, based on these unique identifiers, the corresponding case file association entity identifiers in the source and target case file information are identified. Based on these entity identifiers, the corresponding association pairs are retrieved from the association relationship records to confirm whether an association relationship exists that corresponds to the reference relationship record. If the association relationship shown in the reference relationship record is already recorded in the association relationship record, then the reference relationship record and the association relationship record are considered consistent in that association relationship. If not recorded, it is considered a potentially existing implicit association relationship, and this association relationship is recorded as implicit association mining data. By performing the above verification operations one by one, a comprehensive consistency verification of all reference relationship records is completed, ultimately forming implicit association mining data.

[0040] The implicit association mining data includes association relationship identifiers, corresponding case file association subject identifier pairs, unique identifiers of source case file information and target case file information, implicit association relationship types, and consistency verification results of reference relationship records and association relationship records.

[0041] S6: Update and reconstruct the initial topology of the relationships based on the implicit association mining data to form a full lifecycle association topology of case file information, including: Based on the implicit association mining data, extract the case file association subject identifier pairs corresponding to the implicit association mining data and generate supplementary association relationship records; First, each implicit relationship recorded in the implicit relationship mining data is read one by one. Each implicit relationship includes a relationship identifier, a corresponding case file related subject identifier pair, a unique identifier for the source case file information, a unique identifier for the target case file information, and the type of implicit relationship. For each implicit relationship, the case file related subject identifier pair is extracted. The case file related subject identifier pair consists of two case file related subject identifiers, each corresponding to a unique identifier stored in the case file related subject set, such as personnel name, personnel position or number, full name of unit, event number, etc. The extraction process of the case file related subject identifier pair involves parsing the implicit relationship mining data records one by one, sequentially reading the unique identifiers for the source and target case file information corresponding to the relationship, accurately locating the content of the source and target case file information through the unique identifiers of the source and target case file information, and then using natural language processing technology to extract the corresponding case file related subject identifiers from the source and target case file information respectively. The extracted case file related subject identifier pairs are recorded as structured data of two subject identifiers. Based on the extracted case file related subject identifier pairs, supplementary association relationship records are constructed. The format of the supplementary association relationship records is completely consistent with the defined association relationship records, including the first subject identifier in the related subject pair, the second subject identifier in the related subject pair, the association relationship type, and the time stamps for the first and latest occurrences of the association relationship. The association relationship type is recorded as "implicit association" to distinguish it from the defined co-occurrence relationship and temporal continuity relationship. The time stamp is set as the earliest time of actual occurrence or generation recorded in the source case file information and the target case file information as the first occurrence time, and the latest occurrence time stamp is based on the generation time of the implicit association mining data. For example, if the time stamp of the source case file information is 2024-01-15 09:00:00 and the time stamp of the target case file information is 2024-02-10 15:30:00, then the first occurrence time stamp is set as 2024-01-15 09:00:00, and the latest occurrence time stamp is set as the actual time of generation of the implicit association mining data. All supplementary relationship records are aggregated one by one to form a structured dataset, which records all case file related entity identifiers and related information corresponding to each implicit relationship, thus completing the generation of supplementary relationship records.

[0042] The supplementary relationship records are merged into the relationship records, and the initial relationship topology is updated based on the supplementary relationship records to obtain the updated topology; The generated supplementary relationship records are read one by one and integrated into the already constructed relationship records. For each supplementary relationship record's case file association subject identifier pair, existing relationship records are searched one by one. If the case file association subject identifier pair of the supplementary relationship record is found to be exactly the same as that of an existing relationship record, and the relationship type is the same, then only the latest occurrence time stamp in the existing relationship record is updated to match the supplementary relationship record. If no identical record exists, the supplementary relationship record is directly inserted into the relationship records as a new record. The updated relationship records contain all explicit and implicit relationship information. The initial relationship topology is updated using graph theory, with case file association subject identifiers as topology nodes and relationship records as topology edges. Based on the original initial topology of the association relationships, the updated association relationship records are traversed one by one. For the subject pairs existing in the initial topology of the association relationships, if the association relationship type is consistent and has not changed, the topology edge remains unchanged; if it is a newly added supplementary association relationship record, a corresponding topology edge is added between the topology nodes. The attribute of the newly added topology edge is marked as implicit association, and the first occurrence time and the latest occurrence time mark are recorded, thus obtaining an updated topology that fully reflects the explicit and implicit association relationships.

[0043] In updating the topology, redundant links are eliminated and connectivity checks are performed based on the case file association subject set to generate a reconstructed link set; First, all topological edges in the updated topology are analyzed one by one to identify redundant links. The criteria for determining redundant links are: if two or more topological nodes in the updated topology have the same type, direction, and highly overlapping time markers, they are considered redundant links. The criteria for high overlap are that the overlap ratio between the first and latest occurrence time markers exceeds a set threshold. The threshold is determined by analyzing the statistical characteristics of link time marker overlap in historical data; for example, it can be set to an overlap ratio of 80% or 90%. Redundant links are then eliminated by merging multiple redundant links into a single link. The first occurrence time in the link attribute is set to the earliest first occurrence time among the multiple links, and the latest occurrence time is set to the latest latest occurrence time among the multiple links. After the redundant links are eliminated, a connectivity check is performed to determine whether the connectivity between topological nodes in the updated topology is affected by the elimination of redundant links. The verification method is as follows: traverse all topological nodes in the updated topology structure, using a depth-first search or breadth-first search algorithm to visit all reachable topological nodes one by one, forming a node visit set. If any node is found to be inaccessible from other nodes, the updated topology structure is determined to have insufficient connectivity, and necessary reconstructed links are generated to ensure the updated topology structure has complete connectivity. The reconstructed link set is the complete set of topology links generated or adjusted after redundant link elimination and connectivity verification, recording the corresponding subject identifier pair, association type, and time attribute of each link.

[0044] The updated topology is reconstructed based on the reconstructed link set to obtain the full lifecycle association topology of case file information; Starting with updating the topology, the topological edges between topological nodes are adjusted and supplemented based on the link information recorded in the reconstructed link set. If a link recorded in the reconstructed link set does not exist in the updated topology, a new topological edge is created between the corresponding topological nodes based on the link's subject identifier, recording the link's association type, first appearance time, and latest appearance time marker. If a link already exists but its time marker or type is different, the attributes of the existing topological edge are updated to match the reconstructed link set. After completing the reconstruction or update operations for all reconstructed links, a final association topology reflecting the true state of the case file information's association relationships throughout its entire lifecycle is obtained. The topological nodes of the case file information's lifecycle association topology are the identifiers of all case file association subjects. The topological edges record the association type and time attributes between subjects, and reflect subject association information, including explicit and implicit associations, ensuring accurate representation of the association relationships at each stage of the case file information management lifecycle.

[0045] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0046] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0047] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0048] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0049] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0050] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0051] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0052] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0053] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for managing case file information based on the entire lifecycle, characterized in that, Includes the following steps: S1: Collect case file information generated at different time points throughout the entire life cycle, establish a data index according to the case file information type, and generate a case file information index library; S2: Based on the case file information index, identify the initial relationships between the related entities in the case file information and establish the initial topology of the relationships; S3: Use the link prediction algorithm to analyze the initial topology of the association relationship, identify the weak association regions of the topology that appear over time, and obtain the distribution data of the weak association topology regions. S4: Based on the distribution data of weakly associated topological regions, identify the nodes at risk of breakage in the initial topological structure of the association relationship, and generate a list of predicted degradation of the association of case file related subjects. S5: Based on the node association degradation prediction list, perform association supplement analysis on the nodes at risk of breakage, identify potential implicit associations between case file information, and generate implicit association mining data; S6: Update and reconstruct the initial topology of the association based on the implicit association mining data to form the full life cycle association topology of the case file information.

2. The case file information management method based on the entire lifecycle as described in claim 1, characterized in that, S1, specifically: Collect case file information generated at different time points throughout the entire life cycle and write it into the original case file information set. Standardize the format and unify the character encoding of the case file information and add time stamps and source stamps. Extract case file information type tags based on case file information type mapping rules; Based on the case file information type tags, time markers, and source markers, construct index keys and generate index entries; The aggregated index entries form a case file information index database.

3. The case file information management method based on the entire lifecycle as described in claim 2, characterized in that, S2, specifically: Based on the case file information index database, index entries are read and index keys are parsed to obtain case file information type tags, time stamps, and source stamps; Extract the identifiers of the related entities from the case file information and normalize them to form a set of related entities for the case file. Generate association records based on the co-occurrence and temporal continuity of related subjects in different index entries within the same case file; Aggregate relationship records to construct the initial topology of the relationships.

4. The case file information management method based on the entire lifecycle as described in claim 3, characterized in that, S3, specifically: Based on the initial topology structure of the association relationship, the case file association subject set and association relationship record are read, and the time window is divided according to the time mark and the time window topology structure is generated. In the time window topology, the link prediction algorithm is used to calculate the link prediction score of the case file related subject pairs, forming a link prediction score set; Based on the link prediction score set and the actual link status in the time window topology, the link deviation value is calculated to generate link deviation feature data; Based on the link deviation feature data, weakly correlated node clusters are aggregated according to the weak correlation determination rules, and weakly correlated topological region distribution data is output.

5. The case file information management method based on the entire lifecycle as described in claim 4, characterized in that, S4, specifically: Based on the distribution data of weakly associated topological regions, weakly associated node clusters are read and mapped to the initial topological structure of the association relationship, and the case file association subject set corresponding to the weakly associated node clusters is extracted; In the initial topological structure of the association relationship, the node degree value and intermediate value of the case file association subject set are calculated to form node structure feature data; Calculate node breakage risk values ​​and generate a set of nodes at breakage risk based on node structure feature data and link deviation feature data; Based on the set of fracture risk nodes, output a list of predicted degradation of the associated entities in the case file.

6. The case file information management method based on the entire lifecycle as described in claim 5, characterized in that, S5, specifically: Based on the case file association subject association degradation prediction list, read the set of break risk nodes and locate the index entries corresponding to the set of break risk nodes, and extract the case file information associated with the index entries to form a candidate case file information set; In the candidate case file information set, perform similar merging based on case file information type tags and perform time sorting based on time markers to generate a candidate case file information sequence; Perform keyword matching and citation relationship parsing on the candidate case file information sequence to generate citation relationship records; Consistency checks are performed on reference records and association records to output implicit association mining data.

7. A case file information management method based on the entire lifecycle as described in claim 6, characterized in that, S6, specifically: Based on the implicit association mining data, extract the case file association subject identifier pairs corresponding to the implicit association mining data and generate supplementary association relationship records; The supplementary relationship records are merged into the relationship records, and the initial relationship topology is updated based on the supplementary relationship records to obtain the updated topology; In updating the topology, redundant links are eliminated and connectivity checks are performed based on the case file association subject set to generate a reconstructed link set; The updated topology is reconstructed based on the reconstructed link set to obtain the full lifecycle association topology of case file information.