Case knowledge graph construction method, electronic equipment and storage medium

By constructing a case knowledge graph and using an information extraction model to preprocess case file texts and map entity information, the problem of scattered related case information was solved, achieving efficient structured processing and intelligent analysis of case information, and improving case handling efficiency.

CN122019783APending Publication Date: 2026-05-12BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2025-11-13
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Information on related cases is scattered and inefficient to mine. Existing digital tools lack the ability to perform deep semantic analysis and dynamic relationship mining, resulting in long processing cycles for complex cases and overloading of resources.

Method used

By constructing a case knowledge graph, the case file text is preprocessed using an information extraction model to extract entity information, map it to graph nodes, and construct relation edges. By combining natural language processing and knowledge graph, the structured extraction of case information and the modeling of relationships between entities can be achieved.

Benefits of technology

It enables efficient and structured processing of case information, supports real-time tracking of upstream and downstream cases, improves the intelligence and efficiency of case handling, assists in identifying potential suspicious networks, and provides extended clues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019783A_ABST
    Figure CN122019783A_ABST
Patent Text Reader

Abstract

The invention provides a method for constructing a case knowledge graph, electronic equipment and a storage medium, and relates to the technical field of big data and artificial intelligence application, and the method comprises the steps: obtaining a case file text, and carrying out the preprocessing of the case file text. And extracting the preprocessed case file text through an information extraction model to obtain entity information. And mapping based on the entity information to obtain a graph node, and constructing a relation edge according to the graph node and information in the case file text. And constructing the case knowledge graph based on the graph nodes and the relation edges. According to the method, a framework combining natural language processing and a knowledge graph is introduced, structured extraction of entity information in a case file text is realized, and a graph database is constructed to carry out relationship modeling among entities. The mode breaks through the limitation of a traditional information island, and provides a new thought for upstream and downstream role chain identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of big data and artificial intelligence application technology, and in particular to a method for constructing a case knowledge graph, an electronic device, and a storage medium. Background Technology

[0002] The handling of related cases suffers from problems such as fragmented information and low efficiency in link discovery. Related cases lack structured integration, with upstream and downstream leads scattered across different sectors, and grassroots operations still rely on a rudimentary model of "manual case file review + basic database retrieval." Link discovery is inefficient; existing digital tools are limited to matching single information elements and lack deep semantic analysis and dynamic relationship mining capabilities, resulting in complex cases taking 6-12 months to process and resources operating at a chronically overloaded level. Summary of the Invention

[0003] In view of this, the purpose of this application is to propose a method for constructing a case knowledge graph, an electronic device and a storage medium to solve the problem of low efficiency in mining related cases.

[0004] To achieve the above objectives, the first aspect of this application provides a method for constructing a case knowledge graph, comprising:

[0005] Obtain the case file text and preprocess the case file text; Entity information is obtained by extracting preprocessed case file text using an information extraction model; Graph nodes are obtained based on the entity information mapping, and relation edges are constructed based on the graph nodes and information in the case file text. A case knowledge graph is constructed based on the graph nodes and the relation edges.

[0006] Optionally, the preprocessing of the case file text includes: The case file text was cleaned and formatted. Remove garbled characters from the case file text and then perform unified encoding on the case file text; The case file text is segmented into paragraphs and key fields are indexed.

[0007] Optionally, the step of extracting entity information from the preprocessed case file text using an information extraction model includes: Identify personnel identity information and case information in preprocessed case file texts; Extract contact numbers from preprocessed case file text using regular expressions; Retrieve the financial account information from the pre-processed case file and verify the authenticity of the financial account information; Identify legal entity information in pre-processed case file texts; The identity information, case information, communication account, financial account, and legal entity information are considered as the entity information.

[0008] Optionally, the step of mapping graph nodes based on the entity information includes: A suspicious person node is constructed based on the identity information and the case information; Suspicious account nodes are constructed based on the communication account, the financial account, and other financial information in the pre-processed case file text. Construct a corporate entity node based on the aforementioned legal entity information; The suspicious person node, the suspicious account node, and the corporate legal representative node are used as the graph nodes.

[0009] Optionally, constructing relation edges based on the graph nodes and information in the case file text includes: Based on the information in the case file text, construct graph nodes with edges representing fund transfer relationships, communication relationships, and personnel affiliation relationships. The attributes of the fund transfer relationship edges include at least transaction time, the attributes of the communication relationship edges include at least communication duration, and the attributes of the personnel affiliation relationships edges include at least job title.

[0010] Optional, also includes: The weight of the relation edge is calculated using the following formula: Weight value = ln(transaction amount) × communication frequency coefficient.

[0011] Optional, also includes: In response to obtaining a new case file text, extract the new communication account and new financial account from the new case file text; Based on the new communication account and the new financial account, a match is made in the case knowledge graph. In response to the matching of related cases, the flow of funds between related cases is analyzed and calculated.

[0012] Optionally, it also includes: visualizing the case knowledge graph.

[0013] Based on the same inventive concept, a second aspect of this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.

[0014] Based on the same inventive concept, a third aspect of this application also provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method described above.

[0015] As described above, the case knowledge graph construction method, electronic device, and storage medium provided in this application include: acquiring case file text and preprocessing the case file text; extracting entity information from the preprocessed case file text using an information extraction model; mapping graph nodes based on the entity information; constructing relational edges based on the graph nodes and information in the case file text, providing a data foundation for subsequent case knowledge graph construction; and constructing the case knowledge graph based on the graph nodes and relational edges. This application introduces an architecture combining natural language processing and knowledge graphs to achieve structured extraction of entity information from case file texts and construct a graph database for modeling relationships between entities. This approach breaks the limitations of traditional information silos and provides a new approach for identifying upstream and downstream role chains. Furthermore, based on the constructed case knowledge graph, the linkage between matching algorithms and graph algorithms can assist in identifying potential suspicious networks and provide users with extended clues, demonstrating certain application exploration value. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating the case knowledge graph construction method according to an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the case knowledge graph construction device according to an embodiment of this application; Figure 3 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0020] As described in the background section, current handling of related cases suffers from problems such as fragmented information and low efficiency in link discovery. Currently, related cases lack structured integration, upstream and downstream leads are fragmented, and grassroots operations still rely on a rudimentary model of "manual case file review + basic database retrieval," resulting in inefficient link discovery. Existing digital tools are limited to matching single information elements and lack deep semantic analysis and dynamic relationship mining capabilities, leading to complex cases taking 6-12 months to process and resources operating at a chronic overload.

[0021] Focusing on addressing the issues of information fragmentation and low efficiency in link mining in handling related cases, this application proposes a method for constructing a case knowledge graph. This method features automatic information extraction, graph construction, and intelligent matching, enabling real-time tracking of upstream and downstream cases and providing more intelligent and efficient support for case handling.

[0022] The embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0023] This application proposes a method for constructing a case knowledge graph, such as Figure 1 As shown, it includes the following steps: Step 102: Obtain the case file text and preprocess the case file text.

[0024] Specifically, this embodiment establishes a standardized API interface to connect with the system's internal case management system. When a user submits or updates a case file, the system automatically and in real-time retrieves the case file text via the API. The case file text may contain plain text files, Word documents, PDF files, or even image text that has undergone preliminary OCR processing. In terms of document type, the case file text includes various documents and evidentiary materials, such as non-standard evidentiary materials. These non-standard evidentiary materials include bank statements, chat screenshots, contracts, invoices, company registration information, and explanatory documents issued by third-party institutions. Due to the diversity of data sources, the case file text suffers from problems such as garbled characters, inconsistent encoding, and chaotic formatting.

[0025] Furthermore, the case file text is preprocessed, including: The case file text was cleaned and formatted. Remove garbled characters from the case file text and then perform unified encoding on the case file text; The case file text is segmented into paragraphs and key fields are indexed.

[0026] Specifically, case file texts often contain noisy data, inconsistent formats, and other inconsistencies. Without data cleaning and format standardization, subsequent issues such as entity recognition errors, failed relation extraction, and inaccurate knowledge graph construction can occur. Data cleaning identifies and corrects (or removes) errors, inconsistencies, and irrelevant content from textual data. Format standardization transforms diverse representations of text into a consistent, standardized form. Together, data cleaning and format standardization ensure data quality and consistency, laying a solid foundation for subsequent entity recognition, relation extraction, and other analyses.

[0027] Data cleaning specifically includes removing garbled characters and special characters, removing irrelevant symbols and spaces, and removing irrelevant text headers and footers. Standardizing formats specifically includes standardizing date formats, personal / organization names, and data and amounts.

[0028] When removing garbled characters, predefined regular expression rules are used to match and remove unrecognizable character sequences, special control characters, and "question marks" or "garbled blocks" caused by transmission errors. An encoding tool is used to automatically detect the character encoding of the input text. All text is uniformly converted to a standard encoding (such as UTF-8) to ensure that Chinese characters and other characters can be displayed correctly.

[0029] When segmenting case file text and indexing key fields, the system first automatically identifies the text type (e.g., "bank statements," "chat logs," "business information") based on filename, content keywords, or metadata. For chronological text (such as bank statements), it segments by line breaks to ensure each transaction record is a separate text unit. For question-and-answer text, it segments by speaker identifier, treating each round of questions and answers as a paragraph. For long narrative text, it segments by natural paragraphs (e.g., two line breaks).

[0030] The system assigns predefined tags to different paragraphs or blocks based on the identified document type, serving as indexes for key fields. After preprocessing the case file text, the resulting data is structured / semi-structured text data that is segmented, categorized, and indexed.

[0031] Step 104: Extract entity information from the preprocessed case file text using an information extraction model.

[0032] Specifically, the information extraction model in this embodiment can be a UIE (Universal Information Extraction) model, which leverages its strong generalization ability to identify common and variable entities (such as names and locations). UIE is a general-purpose information extraction model. A major advantage is its ability to quickly adapt to various extraction tasks (such as entity recognition, relation extraction, and event extraction) with structured pattern prompts, requiring little to no training. For example, the information extraction model might be prompted to find entities such as "find all persons in the text," "identify all crime scenes," and "extract all time points." The model then outputs a structured recognition result: {"Entity": "Zhang San", "Type": "Case Personnel", "Starting Position": 105, "Confidence": 0.98}.

[0033] Furthermore, the entity information obtained by extracting entity information from the preprocessed case file text using the information extraction model includes: Identify personnel identity information and case information in preprocessed case file texts; Extract contact numbers from preprocessed case file text using regular expressions; Retrieve the financial account information from the pre-processed case file and verify the authenticity of the financial account information; Identify legal entity information in pre-processed case file texts; The identity information, case information, communication account, financial account, and legal entity information are considered as the entity information.

[0034] Specifically, personnel identity information and case information are extracted using an information extraction model. Personnel identity information includes the person's name. Case information includes the location of the incident and key time points. Key time points include the time of the incident, the time of the transaction, etc.

[0035] This method extracts contact information (such as WeChat IDs) from case file text using regular expressions. Since contact information typically follows a relatively fixed pattern, regular expressions can accurately extract it. The process involves writing a regular expression pattern to match contact information. For example, a regular expression (e.g., ` / wxid_[a-z0-9]{8,} / `) is used to extract WeChat IDs, filtering out non-standard characters. After matching, a second filtering step is performed to remove characters that clearly do not conform to the pattern of contact information, retaining only compliant matches.

[0036] Financial account numbers can also be extracted from case file text using regular expressions. These accounts are typically consecutive digits, such as bank account numbers which are 16-19 consecutive digits. To ensure the accuracy of the extracted financial account numbers, their authenticity needs to be verified. For example, the Bank Identification Number (BIN) can be extracted from the bank account number and sent to a BIN query interface provided by the bank or UnionPay. If the interface returns information such as the bank name and card type corresponding to the BIN number, the verification is successful. If the interface returns "BIN number does not exist," the verification fails, and the financial account number is discarded.

[0037] When identifying legal entity information, sentence structure rules and entity association can be used. Legal entity information includes the legal entity's name, company name, and legal representative. Sentence structure rules or a small NLP model can be used to identify sentences containing keywords such as "legal representative" or "legal person's representative." For example, the sentence structure could be "Legal Representative: {Name}". The {Name} part can be extracted from this sentence structure. The legal entity's name is then associated with the "company name" and "business registration code" in the context to form a complete triple, such as <Company Name, Legal Representative, Legal Entity Name>.

[0038] The entity information extraction method in this step transforms unstructured dossier text into high-quality, structured data, providing a crucial information foundation for subsequent knowledge graph construction.

[0039] Step 106: Obtain graph nodes based on the entity information mapping, and construct relation edges based on the graph nodes and information in the case file text.

[0040] Specifically, structured entity information is mapped to nodes in a graph; for example, entity data is mapped to nodes in the graph database Neo4j. Each node has labels and attributes. Meaningful connections, or relationship edges, are established between related nodes based on information in the case file text. For example, in Neo4j, relationships are established using the CREATE or MERGE statements of the Cypher query language.

[0041] Furthermore, the step of obtaining graph nodes based on the entity information mapping includes: A suspicious person node is constructed based on the identity information and the case information; Suspicious account nodes are constructed based on the communication account, the financial account, and other financial information in the pre-processed case file text. Construct a corporate entity node based on the aforementioned legal entity information; The suspicious person node, the suspicious account node, and the corporate legal representative node are used as the graph nodes.

[0042] The graph nodes include suspicious person nodes, suspicious account nodes, and corporate entity nodes. Suspicious person nodes are constructed based on the identity information and case information. Attributes of suspicious person nodes may include name, ID number, suspicious events, etc. Suspicious account nodes are constructed based on the communication account, the financial account, and other financial information in the pre-processed case file text. Attributes of suspicious account nodes include WeChat ID or bank card number, bank name, and account balance, etc. Corporate entity nodes are constructed based on the corporate entity information. Attributes of corporate entity nodes include company name, business registration code, and industry type, etc.

[0043] Furthermore, the step of constructing relational edges based on the graph nodes and information in the case file text includes: Based on the information in the case file text, construct graph nodes with edges representing fund transfer relationships, communication relationships, and personnel affiliation relationships. The attributes of the fund transfer relationship edges include at least transaction time, the attributes of the communication relationship edges include at least communication duration, and the attributes of the personnel affiliation relationships edges include at least job title.

[0044] Specifically, relationship edges include those related to fund transfers, communication connections, and personnel affiliations. Relationship edge construction is an evidence-driven, rule-oriented, automated process. It's not simply about connecting nodes; rather, it transforms information from each piece of case evidence into concrete, attribute-laden connecting lines in a graph. The system scans the information in the case file text, automatically determining the type of relationship described by the evidence based on keywords and context, and triggering the corresponding relationship construction rules.

[0045] Fund flow relationship edges describe the path of funds flowing between entities and can be established using bank transaction records, transfer vouchers, transfer information mentioned in chat logs, and descriptions of fund transactions. Date and time information, transaction amounts, and bank transaction numbers are extracted from the text and used as attributes of the fund flow relationship edges.

[0046] Communication relationship edges describe the social network and the closeness of connections between entities. They can be established using call logs, WeChat chat history, and descriptions of interactions between the parties. The aggregated number of calls, call or contact time, and cumulative call duration are used as attributes of the communication relationship edges.

[0047] Personnel affiliation edges describe the legal or employment relationship between an individual and an organization (company). This can be established through business registration information, company bylaws, and statements about job positions in employment contracts. Job role and start date of employment are attributes of personnel affiliation edges.

[0048] Furthermore, after constructing the relationship edges, it is also necessary to calculate the weight value of each relationship edge. The weight value of the relationship edge is calculated using the following formula: Weight value = ln(transaction amount) × communication frequency coefficient.

[0049] Specifically, when a new relationship edge is created, an operation to calculate the weight value for that relationship edge is triggered. The weight value quantifies the tightness of the case association. The system reads the attributes of the relationship edge and its related edges, extracting the transaction amount and the total number of communications. A communication frequency coefficient is calculated based on the total number of communications. For example, the communication frequency coefficient C can be calculated as: C = 1 + α × ln(1 + total number of communications N), where α is an adjustable amplification coefficient that determines the contribution of the total number of communications to the final weight. When there is no communication (N=0): coefficient C = 1 + α × ln(1+0) = 1. At this time, the weight value = ln(transaction amount) × 1, and the weight is entirely determined by the transaction amount. When there is communication (N>0): coefficient C>1. The more total communications, the larger the coefficient C, thus proportionally amplifying the basic association weight established by the financial transactions.

[0050] By introducing a logarithmic function and an adjustable amplification factor in the calculation of weight values, the vastly different original data (amount, number of times) are normalized to a comparable scale. Furthermore, by integrating the two key dimensions of "fund flow" and "information flow", a quantitative indicator that can relatively scientifically reflect the "closeness of case correlation" is finally obtained.

[0051] Relationship edge construction is a refined, evidence-based mapping process. Relationship edges transform actions recorded in case file texts (transfers, phone calls, appointments) into semantically rich, quantifiable connections within a knowledge graph. This lays a solid data foundation for subsequent real-time collision analysis, path analysis, and visual early warning systems.

[0052] Step 108: Construct a case knowledge graph based on the graph nodes and the relation edges.

[0053] Specifically, after constructing the graph nodes and relation edges, the case knowledge graph is built by connecting the graph nodes and relation edges. The case knowledge graph enables structured storage of all case entities and relationships, dynamically maintaining the accuracy and timeliness of the graph data. It provides a data foundation for intelligent identification of complex case chains and patterns, as well as real-time collision detection and risk warning.

[0054] Based on steps 102 to 108 above, the method for constructing a case knowledge graph provided in this embodiment includes obtaining case file text and preprocessing the case file text. Entity information is extracted from the preprocessed case file text using an information extraction model. Graph nodes are mapped based on the entity information, and relational edges are constructed according to the graph nodes and information in the case file text, providing a data foundation for subsequent case knowledge graph construction. The case knowledge graph is constructed based on the graph nodes and relational edges. This application introduces an architecture combining natural language processing and knowledge graphs to achieve structured extraction of entity information from case file texts and to construct a graph database for modeling relationships between entities. This approach breaks the limitations of traditional information silos and provides a new approach for identifying upstream and downstream role chains. Furthermore, based on the constructed case knowledge graph, the linkage between matching algorithms and graph algorithms can assist in identifying potential suspicious networks and provide users with extended clues, demonstrating certain application exploration value.

[0055] In some embodiments, it also includes: In response to obtaining a new case file text, extract the new communication account and new financial account from the new case file text; Based on the new communication account and the new financial account, a match is made in the case knowledge graph. In response to the matching of related cases, the flow of funds between related cases is analyzed and calculated.

[0056] Specifically, when new case file texts are available, it is necessary to dynamically discover deep-seated, cross-case relationships and risks from a static knowledge graph. The system may simultaneously input multiple new case file texts, requiring the system to handle a large number of query requests concurrently. To achieve sub-second response times, high-concurrency technology is essential. Therefore, this embodiment establishes a lightweight thread pool to support concurrent processing of real-time collision requests. In practice, a batch of processing threads (e.g., 100) is pre-created in the server background. When a new case file text is available, an idle thread is directly selected from the pool for processing, avoiding the significant overhead of frequently creating and destroying threads. The thread pool ensures that the system can handle more than 1000 collision requests simultaneously, guaranteeing the system's response speed and stability under high load. Simultaneously, an inverted index table is established based on a Redis in-memory database. When a new case file text is entered, its communication account or financial account is scanned in real time, achieving sub-second collision matching with historical case file texts. By establishing the inverted index table, a list of all case IDs that have appeared can be retrieved based on the communication account or financial account. When a new case file is entered, its unique new communication account and financial account are extracted. The system then concurrently queries Redis. Redis stores all data in memory, with extremely fast read and write speeds (up to 100,000 times / second or more), and query operations can be completed in milliseconds. If Redis returns a non-empty result (e.g., bankcard:6225880100123456 associated with [case_033, case_078]), the system immediately knows that the new case file is associated with historical case files 033 and 078. Historical case files 033 and 078 are the associated cases of the new case file, thus achieving a second-level collision detection.

[0057] When a communication account appears in three or more different cases, the system automatically identifies it as a crucial connecting thread. From the complete case knowledge graph, the system extracts all graph nodes and edge relationships involved in all cases containing this communication account, forming a smaller, more focused subgraph for in-depth analysis. Using Dijkstra's algorithm on this subgraph, starting from a key account, the system explores all possible paths leading to other accounts, calculating the cumulative weight of each path. For edge weights, higher weights indicate a more important funding or relationship chain. Paths with a cumulative weight >100 are designated as high-risk paths. These high-risk paths are marked with a special color, such as red, in the graph. The algorithm automatically identifies accounts or individuals appearing in numerous high-risk paths; these are key funding hubs. Finally, an analysis report containing a list of high-risk paths and funding hub nodes is generated to directly assist user decision-making. The method in this embodiment solves the problem of information silos, achieves second-level association and matching between new cases and massive historical case data, and uses graph algorithms to automatically identify the most dangerous and critical paths, dynamically discovering deep-seated cross-case associations and risks from static knowledge graphs.

[0058] In some embodiments, the method further includes: visualizing the case knowledge graph.

[0059] Specifically, by visualizing the case knowledge graph, it is transformed into an easily understandable visual language, ensuring the security and compliance of data access. A circular, focused layout visualization engine is implemented using the Vue.js + D3.js technology stack, presenting the case network relationship diagram in a radial format radiating from a central node. Vue.js is a front-end framework responsible for building the user interface of the entire web application, handling user interactions, and managing application state, making development more modular and efficient. D3.js is a powerful data visualization library used to manipulate the Document Object Model, generate complex charts based on data, and handle calculations of the position of each node in the graph, drawing connections, and processing dragging and zooming. The central node is the key person or account currently being investigated by the user. The first ring in the radial circle consists of entities directly related to the central node (called first-degree relationships), the second ring consists of entities directly related to entities in the first ring (called second-degree relationships), and so on.

[0060] An integrated dynamic data masking gateway implements differentiated information display based on user roles and permissions. Users responsible for handling cases are shown complete information about that case, while other information is masked. Typically, user systems have strict confidentiality requirements and data access restrictions. A user usually does not have permission to view all sensitive information (such as ID number, complete bank account number) for cases they haven't handled. The dynamic data masking gateway acts as a security layer between the front-end and back-end data. When a user logs into the system, the system identifies their identity and permissions. The user requests to view a graph, and the back-end sends the complete graph data to the dynamic data masking gateway. The gateway filters the data according to the user's permission list. Data for cases handled by the user is displayed in full; for data from cases not handled by the user, sensitive fields are masked, including communication accounts, identity information, and financial accounts. The masked, secure data is then sent to the front-end for display to the user. This ensures visibility while protecting personal privacy and case confidentiality.

[0061] In the initial display interface, only the core nodes and their first-degree relationships are shown. When a user clicks on a node with a first-degree relationship, the click event is captured via Vue.js, and more relationship data for that node is requested from the backend. Using dynamic effects from D3.js, the upstream and downstream entities of that node are smoothly expanded and displayed, achieving a layered, in-depth display. This display method enables the case chain drill-down function. In the front-end code, each edge (relationship) in the graph is evaluated. If the weight of an edge or the cumulative weight of its path is greater than 100, the edge flashes red to attract the user's attention and help them quickly locate and view key information. This method allows complex graph data and algorithm results to be presented to the user in an intuitive, secure, and interactive manner.

[0062] This application utilizes multimodal entity intelligent extraction, dynamic graph weight modeling, and cross-file second-level collision technology to achieve structured extraction of elements such as communication accounts, financial accounts, and legal entity information from case materials. This automatically constructs a case association network and provides quantitative early warnings for high-risk paths. Combined with a permission-aware visualization engine, it solves the problem of information silos in cases and improves the efficiency of case chain analysis. This application can be effectively applied to information mining of upstream and downstream economic cases, helping the system quickly understand case relationships and trends from multiple dimensions and improving the level of information management.

[0063] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.

[0064] It should be noted that some embodiments of this application have been described above. In some cases, the actions or steps described in the above embodiments can be performed in a different order than that shown in the above embodiments and the desired result can still be achieved. In addition, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0065] Based on the same inventive concept, corresponding to any of the above embodiments, this application also provides a case knowledge graph construction device.

[0066] refer to Figure 2 The case knowledge graph construction device includes: The acquisition module 202 is configured to acquire case file text and preprocess the case file text. Extraction module 204 is configured to extract entity information from preprocessed case file text using an information extraction model; The mapping module 206 is configured to obtain graph nodes based on the entity information and construct relation edges based on the graph nodes and information in the case file text. The construction module 208 is configured to construct a case knowledge graph based on the graph nodes and the relation edges.

[0067] In some embodiments, the acquisition module 202 is further configured to clean and standardize the format of the case file text; remove garbled characters from the case file text and uniformly encode the case file text; and perform paragraph segmentation and key field indexing on non-standard files in the case file text.

[0068] In some embodiments, the extraction module 204 is further configured to: identify identity information and case information in the preprocessed case file text; extract communication accounts in the preprocessed case file text using regular expressions; retrieve financial accounts in the preprocessed case file text and verify the authenticity of the financial accounts; identify legal person information in the preprocessed case file text; and use the identity information, the case information, the communication accounts, the financial number combination, and the legal person information as the entity information.

[0069] In some embodiments, the mapping module 206 is further configured to construct a suspicious person node based on the identity information and the case information; construct a suspicious account node based on the communication account, the financial account, and other financial information in the preprocessed case file text; construct a corporate legal person node based on the legal person information; and use the suspicious person node, the suspicious account node, and the corporate legal person node as the graph nodes.

[0070] In some embodiments, the mapping module 206 is further configured to construct fund transfer relationship edges, communication relationship edges, and personnel affiliation relationship edges between graph nodes based on information in the case file text. The attributes of the fund transfer relationship edges include at least transaction time, the attributes of the communication relationship edges include at least communication duration, and the attributes of the personnel affiliation relationship edges include at least job title.

[0071] In some embodiments, a calculation module is also included, configured to calculate the weight value of the relation edge using the following formula: weight value = ln(transaction amount) × communication frequency coefficient.

[0072] In some embodiments, the matching module is configured to, in response to obtaining new case file text, extract new communication accounts and new financial accounts from the new case file text; perform matching in the case knowledge graph based on the new communication accounts and new financial accounts; and, in response to matching related cases, analyze and calculate the flow of funds between related cases.

[0073] In some embodiments, a visualization module is also included, configured to visualize the case knowledge graph.

[0074] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.

[0075] The apparatus described above is used to implement the corresponding case knowledge graph construction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0076] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the case knowledge graph construction method described in any of the above embodiments.

[0077] Figure 3This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0078] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0079] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0080] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0081] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0082] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0083] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0084] The electronic devices described above are used to implement the corresponding case knowledge graph construction methods in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0085] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to execute the case knowledge graph construction method as described in any of the above embodiments.

[0086] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0087] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the case knowledge graph construction method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0088] Based on the same concept, corresponding to any of the above embodiments, this application also provides a computer program product, including computer program instructions, which, when run on a computer, cause the computer to perform the method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0089] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.

[0090] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to choose, based on the prompt message, whether to provide personal information to the software or hardware such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution.

[0091] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" regarding the provision of personal information by the electronic device.

[0092] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0093] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application is limited to these examples; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in detail for the sake of brevity.

[0094] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0095] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0096] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.

Claims

1. A method for constructing a case knowledge graph, characterized in that, include: Obtain the case file text and preprocess the case file text; Entity information is obtained by extracting preprocessed case file text using an information extraction model; Graph nodes are obtained based on the entity information mapping, and relation edges are constructed based on the graph nodes and information in the case file text. A case knowledge graph is constructed based on the graph nodes and the relation edges.

2. The method according to claim 1, characterized in that, The preprocessing of the case file text includes: The case file text was cleaned and formatted. Remove garbled characters from the case file text and then perform unified encoding on the case file text; The case file text is segmented into paragraphs and key fields are indexed.

3. The method according to claim 1, characterized in that, The process of extracting entity information from the preprocessed case file text using an information extraction model includes: Identify personnel identity information and case information in preprocessed case file texts; Extract contact numbers from preprocessed case file text using regular expressions; Retrieve the financial account information from the pre-processed case file and verify the authenticity of the financial account information; Identify legal entity information in pre-processed case file texts; The identity information, case information, communication account, financial account, and legal entity information are considered as the entity information.

4. The method according to claim 3, characterized in that, The process of mapping graph nodes based on the entity information includes: A suspicious person node is constructed based on the identity information and the case information; Suspicious account nodes are constructed based on the communication account, the financial account, and other financial information in the pre-processed case file text. Construct a corporate entity node based on the aforementioned legal entity information; The suspicious person node, the suspicious account node, and the corporate legal representative node are used as the graph nodes.

5. The method according to claim 1, characterized in that, The construction of relation edges based on the graph nodes and information in the case file text includes: Based on the information in the case file text, construct graph nodes with edges representing fund transfer relationships, communication relationships, and personnel affiliation relationships. The attributes of the fund transfer relationship edges include at least transaction time, the attributes of the communication relationship edges include at least communication duration, and the attributes of the personnel affiliation relationships edges include at least job title.

6. The method according to claim 1, characterized in that, Also includes: The weight of the relation edge is calculated using the following formula: Weight value = ln(transaction amount) × communication frequency coefficient.

7. The method according to claim 1, characterized in that, Also includes: In response to obtaining a new case file text, extract the new communication account and new financial account from the new case file text; Based on the new communication account and the new financial account, a match is made in the case knowledge graph. In response to the matching of related cases, the flow of funds between related cases is analyzed and calculated.

8. The method according to claim 1, characterized in that, Also includes: The knowledge graph of the case is visualized.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 8.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 8.