Self-adaptive knowledge graph construction and deep analysis method supporting multiple types of files
Through the adaptive knowledge graph construction method, file entities are automatically parsed, cross-file association relationships are established, and processing matters and templates are matched, which solves the problem of inefficiency of traditional manual processing and realizes efficient and accurate file processing and intelligent services.
Patent Information
- Application Number
- CN202510574039.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-06
AI Technical Summary
Traditional manual file processing methods are inefficient and error-prone when facing a large number of documents, and cannot meet the needs of multi-document review, especially in the fields of law, finance, scientific research and government affairs.
Adaptive knowledge graph construction method is adopted to analyze file entities, establish cross-file association relationships, match processing matters and templates, automatically execute processing content using the processing engine, and predict derivative needs, resolve conflict problems, and optimize processing flow.
It realizes efficient and accurate file processing, improves processing efficiency, reduces manual errors, and provides intelligent service upgrades and user experience optimization.
Smart Images

Figure CN120409649A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of document processing, and in particular, to an adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of documents. Background Art
[0002] With the development of informatization, the number of document materials that enterprises and individuals need to process and review has increased exponentially. Especially in the fields of law, finance, scientific research, and government affairs, the need for multi-document review has become increasingly prominent. The traditional document processing and review methods mainly rely on manual viewing and comparison of document content. When there are a large number of document materials and numerous review points that need to be viewed and compared, the above-mentioned manual processing method is not only inefficient but also prone to incorrect reviews or missed review points due to human factors, so it needs to be improved. Summary of the Invention
[0003] In order to achieve high-efficiency and high-accuracy processing of multiple document materials, this application provides an adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of documents.
[0004] In a first aspect, this application provides an adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of documents, adopting the following technical solutions: Receive a processing instruction proposed by a user, and parse all the documents included in the processing instruction. The parsing operation at least includes extracting document entities; Based on the parsing result, match a processing item and a processing template corresponding to the processing item for the processing instruction; wherein, the processing template contains processing content for processing entities in the document; Based on a preset entity alignment technology, determine and establish an association relationship among all the documents included in the processing instruction; According to the determined processing template and the association relationship among the documents, execute all the processing content through a preset processing engine, and feedback the processing result to the user.
[0005] By adopting the above technical solutions, for multiple document materials input by the user, the present application uses the above technical solutions to automatically parse and extract key content (such as entities) in each document, and match specific processing matters based on the key content of the document, so as to adaptively learn the user's processing requirements by analyzing the document content without knowing the specific processing purpose of the processing instruction triggered by the user; then, based on the entity alignment technology and the document entities that have been extracted previously, establish cross-document association relationships, and use the document entities as the link between the association relationships and the processing content to achieve the mutual correspondence between the association relationships and the processing content, so that when the processing engine executes the processing content, it can quickly lock the documents to be processed and the specific entities in the documents according to this corresponding relationship, so as to efficiently summarize the association relationships between documents from the complex multiple documents, help improve the processing efficiency, and use this intelligent processing method to replace manual processing, which can also solve the problem of affecting the processing accuracy due to human errors.
[0006] Optionally, the method further includes: Whenever a processing matter is matched, determine whether there is a derivative requirement that meets the preset derivative condition. If so, determine the derivative matter corresponding to the derivative requirement, and generate and feedback a derivative matter processing guide to the user according to the processing template corresponding to the derivative matter.
[0007] By adopting the above technical solutions, through the processing implementation directly matched, autonomously mine derivative requirements, realize the upgrade from "passive response" to "active service", deeply and long-term analyze and mine user requirements, plan and guide the processing solutions of matters for the user from a long-term perspective, optimize the intelligence and automation level of the service, and improve the user experience.
[0008] Optionally, the determining whether there is a derivative requirement that meets the preset derivative condition. If so, determining the derivative matter corresponding to the derivative requirement includes: Through a pre-constructed matter relationship graph, determine whether there is a relationship chain that includes the processing matter. The relationship chain also includes other matters except the processing matter, and there is an association relationship between adjacent matters in the relationship chain, and there is an association requirement for describing the association relationship; wherein, the matter relationship graph is used to represent the association relationship between different matters; Based on all the documents included in the processing instruction, calculate the association strength between each matter in the association chain and the processing matter respectively, and take the matter whose association strength meets the preset derivative condition as the derivative matter, and take the corresponding association requirement as the derivative requirement; wherein, the association strength at least includes: the processing completion rate of using all the documents included in the processing instruction to process the matter to which it belongs.
[0009] By adopting the above technical solution, the present application presets a matter relationship graph for describing the association relationship between different matters, and each matter in the matter relationship graph is bridged through association requirements. Based on this relationship graph, all matters associated with the currently processed matter are captured, thereby realizing the leap from single-point matching to global matter intelligent mining, helping to more comprehensively mine the long-term needs of users. Further, the present application realizes the accurate capture and intelligent screening of derivative matters by calculating the association strength between each matter and the processed matter. When calculating the association relationship, the present application fully considers the processing completion rate of the currently existing documents when executing derivative matters, so as to improve the screening accuracy of derivative matters and make the screened derivative matters more in line with the actual objective conditions.
[0010] Optionally, the parsing operation further includes semantic recognition of the file and extraction of keywords. Based on the parsing result, matching a processing matter and a processing template corresponding to the processing implementation for the processing instruction includes: Based on the data sets corresponding to each matter stored in the preset structured database, comparing the similarity between the keywords in the parsing result and the data sets to determine all alternative matters and the similarity corresponding to each alternative matter. Taking the alternative matter with the highest similarity as the processing matter and determining the corresponding processing template, and taking all alternative matters other than the processing matter as derivative matters.
[0011] By adopting the above technical solution, the present solution specifically discloses a matching logic for matching processing matters. When faced with multiple matters (i.e., alternative matters) similar to the file content (i.e., keywords), the present application proposes to select the matter with the highest similarity as the processing matter, that is, to limit the uniqueness of the processing matter. At the same time, in order to ensure that the matters actually wanted by the user can be completed more efficiently, the present application takes all alternative matters that are not selected as processing matters as derivative matters, so as to connect with the subsequent processing scheme of derivative matters mentioned above.
[0012] Optionally, the method further includes: Determining whether there is a conflict problem between the processing matter and all its corresponding derivative matters. If so, based on the conflict type corresponding to the conflict problem, determining the conflict resolution scheme corresponding to the conflict type, and adding the conflict problem and its corresponding conflict resolution scheme to the derivative matter processing guide. Wherein, the conflict problem refers to the situation where the entities shared by the processing matter and the derivative matter are processed in a contradictory manner according to the corresponding processing content. The conflict types at least include: when the processing matters and the derivative matters execute the corresponding processing contents on the shared entity, contradictions occur due to the inconsistent specific contents of the shared entity.
[0013] By adopting the above technical solution, the present application further analyzes the processing matters and the derivative matters to predict whether conflict problems will occur during the execution process of the processing matters and the derivative matters. The essence of the conflict problem is the contradiction between the specific processing contents included in the corresponding processing templates of each matter. For example, the files uploaded by the user include File A: business license (including address a), File B: new lease contract (including address b), and the finally selected processing matter is "annual inspection of business license", including the derivative matter: "recording of address change"; there is a shared entity: "address". However, when executing the processing matter, it is required to lock the original address (i.e., address a), while when processing the above derivative matter, it is required to overwrite the original address (i.e., address a) with address b. Therefore, conflicting operations occur when processing the processing matter and the derivative matter. In this regard, the present application proposes to perform conflict prediction on the execution process of the processing matter and all its corresponding derivative matters, and propose corresponding conflict resolution schemes to resolve the conflict problem.
[0014] Optionally, the method further includes: Regularly, based on the conflict problems discovered historically, regarding the matters related to the conflict problems as target matters; Disassembling and reorganizing all the processing contents included in the processing template corresponding to the target matter, and satisfying: there are conflict contents in the reorganized processing contents, and the conflict contents are the processing contents only including the corresponding conflict problems, establishing and storing the conflict relationship between the conflict contents, the conflict problems, and the corresponding matters; The conflict resolution scheme at least includes preferentially processing the processing contents other than the conflict contents.
[0015] By adopting the above technical solution, regularly realizing the refinement disassembly and reorganization of the specific processing contents of each matter according to the conflict problems between the matters, so that when a conflict problem occurs, preferentially execute the other processing contents other than the conflict contents included in the processing contents of multiple matters. On the one hand, it can improve the conflict detection granularity, and on the other hand, it can also realize parallel processing of the processing contents without conflict problems between multiple matters, improving the processing progress of each matter.
[0016] Optionally, the method further includes: Whenever a processing matter is matched, determining and storing the matching relationship between the processing matter and its corresponding key file, where the key file refers to the file that plays a decisive role in the process of matching and determining the processing matter; Regularly analyze the sensitivity of each key document based on the matching relationships stored in historical periods, and label the key documents with sensitivity higher than a preset sensitive threshold as highly sensitive documents. Determine a set of matters for each of the highly sensitive documents, where the set of matters contains the matters that have a matching relationship with the corresponding highly sensitive document. When matching processing instructions with processing matters, if there are highly sensitive documents among all the documents included in the processing instructions, preferentially retrieve the matters in the set of matters corresponding to the highly sensitive documents to perform the matching operation.
[0017] By adopting the above technical solution, analyze the relationship between files and processing matters, determine the file sensitivity, and the file sensitivity can be used to characterize the scarcity of files. For example, files that are only used when handling certain special matters (i.e., highly sensitive documents). By labeling highly sensitive documents and using them in subsequent matching operations for processing matters, and by retrieving the matters corresponding to highly sensitive documents for the matching operation, the matching efficiency of processing matters can be improved.
[0018] In a second aspect, the present application provides an adaptive knowledge graph construction and in-depth analysis system that supports multiple types of files, including: A file parsing module for receiving a processing instruction proposed by a user and parsing all the files included in the processing instruction. The parsing operation at least includes extracting file entities. A matter matching module for, based on the parsing result, matching a processing matter and a processing template corresponding to the processing matter for the processing instruction; wherein, the processing template contains processing content for processing entities in the file. A cross-file association module for, based on a preset entity alignment technology, determining and establishing an association relationship among all the files included in the processing instruction. A matter processing module for, according to the determined processing template and the file inter-association relationship, executing all the processing content through a preset processing engine and feeding back a processing result to the user.
[0019] In a third aspect, the present application provides an adaptive knowledge graph construction and in-depth analysis device that supports multiple types of files, including a memory and a processor. A computer program capable of being loaded and executed by the processor, such as the method described in any one of the first aspects, is stored on the memory.
[0020] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program capable of being loaded and executed by the processor, such as the method described in any one of the first aspects.
[0021] In summary, the present application includes at least one of the following beneficial technical effects: 1. In this application, by obtaining multiple documents and achieving automatic parsing of each document, and based on the parsing results, matching specific processing matters, so as to adaptively learn the user's processing requirements by analyzing the document content without knowing the specific processing purpose of the user's triggered processing instruction; 2. Further, based on the knowledge graph data and entity alignment technology, this application establishes cross-document association relationships, and then based on these association relationships and processing templates corresponding to the processing matters, finally realizes the efficient processing of the processing matters. In summary, by efficiently extracting and summarizing the association relationships between documents from a large number of complex documents, it helps to improve the processing efficiency, and replacing manual processing with this intelligent processing method can also solve the problem of affecting the processing accuracy due to human errors. Brief Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0023] Figure 1 It is a schematic flowchart of an adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of documents according to an embodiment of the present application.
[0024] Figure 2 It is a structural block diagram of an adaptive knowledge graph construction and in-depth analysis system for supporting multiple types of documents according to an embodiment of the present application.
[0025] Description of the reference numerals: 201, document parsing module; 202, matter matching module; 203, cross-document association module; 204, matter processing module. Detailed Embodiments
[0026] The following will further describe the present application in detail Figure 1-2 in conjunction with the attached
[0027] An embodiment of the present application discloses an adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of documents (hereinafter simply referred to as the analysis method), aiming to efficiently parse multi-source documents to obtain user requirements, and then realize the response and processing of user requirements. The execution subject of the analysis method is an adaptive knowledge graph construction and in-depth analysis system for supporting multiple types of documents (hereinafter simply referred to as the analysis system). The following will specifically describe the specific execution steps of the analysis system for the analysis method in conjunction with Figure 1 specifically elaborate.
[0028] S101. Receive the processing instructions proposed by the user, and parse all the files included in the processing instructions. The parsing operation at least includes extracting file entities.
[0029] In implementation, exemplarily, the user can access the analysis system in the form of a web page and upload multi-source files to trigger the processing instructions. Multi-source files refer to multiple files, and the file formats (such as PDF, Word / image, etc.) and file contents can be different from each other. After the analysis system receives all the files included in the processing instructions, it will parse the files. The parsing operation specifically includes: using open-source libraries (such as Apache PDFBox, pdfplumber / docx libraries in Python) to extract text and tables in PDF and Word formats, identifying files in images and scanned documents through OCR technologies (such as Tesseract, Alibaba Cloud OCR API), and then extracting entities (such as person names, addresses, dates, etc.) from the content extracted above through pre-trained models (such as SpaCy, Alibaba Cloud NLP).
[0030] S102. Based on the parsing result, match a processing matter and a processing template corresponding to the processing matter; wherein, the processing template contains processing content for processing entities in the file.
[0031] Wherein S102 specifically includes the following sub-steps: Based on the data sets corresponding to each matter stored in the preset structured database, compare the keywords in the parsing result with the data sets to determine all alternative matters and the similarity corresponding to each alternative matter; take the alternative matter with the highest similarity as the processing matter and determine the corresponding processing template.
[0032] In implementation, the processing matter is used to represent the matter that the user needs to handle. The analysis system will analyze the processing matter that the user wants to handle, that is, match the processing matter, according to all the files entered when the user triggers the processing instructions and according to the parsing result of the file content. Correspondingly, the analysis system pre-constructs a structured knowledge base, and the structured knowledge base can specifically be a graph database (Neo4j) or a relational database, to store the specific file content required for each processing matter and realize the corresponding relationship of "processing matter - required file content". For example, when the processing matter is "opening a store", the corresponding files that need to be provided are "business license", "lease contract", "health permit", etc.
[0033] The analysis system compares the specific file content corresponding to all processing matters stored in the knowledge base with the parsing result in turn and calculates the similarity, such as using an NLP model (such as BERT) to calculate the similarity. Exemplarily: if the keywords "lease" and "store" are included in the parsing result, the corresponding processing matter with the matching content is "opening a store".
[0034] If there are multiple matters involved in the above comparison process, such as when "Fire Inspection Checklist" also appears in the analysis result, the processing matter corresponding to the matching content "Hygiene Permit" shall be; in this regard, the present application further proposes to calculate the matching score by combining the TF-IDF keyword weight and the BERT sentence embedding similarity, and select the matter with the highest matching score as the processing matter, that is, select the matter with the highest matching degree as the only processing matter.
[0035] Furthermore, each processing matter also corresponds to a processing template, and the processing template includes the specific processing content for processing the corresponding processing matter. The processing content can be expressed in the form of "entity + processing operation", that is, it specifically describes the specific processing operation performed on the entity. For example, when the processing matter is "opening a store", at least one piece of processing content included in its corresponding processing template is: "Verify the consistency of the business license address and the lease contract address". The entity corresponding to this content is "address", and "business license address" and "lease contract address" are both specific description methods of the "address" entity and come from different documents. In this regard, the present application proposes to use entity alignment technology to establish the association between files to facilitate the processing of the above processing content. The corresponding process steps are as follows: S103, Based on the preset entity alignment technology, determine and establish the association relationship among all the files included in the processing instruction; S104, According to the determined processing template and the file association relationship, execute all the processing content through the preset processing engine, and feedback the processing result to the user.
[0036] In implementation, the analysis system pre-constructs a knowledge graph for each entity. The knowledge graph contains all the description contents for describing the same entity and the file to which each description content belongs. Therefore, the analysis system can first establish the association between different description contents of the same entity through entity alignment technology and the aforementioned knowledge graph, and then establish the file association relationship based on the files to which the description contents belong. Finally, based on this association relationship, execute the processing content in the processing template until the processing of all the processing content in the processing template is completed and the processing result is feedback to the user. The processing result specifically includes processing completed and processing failed. And for the situation where the file association cannot be established due to the lack of files included in the processing instruction, resulting in unprocessable processing content, the corresponding feedback processing result is processing failed. At this time, the analysis system is used to feedback the failure reason, that is, output the specific processing content of the processing failure.
[0037] Optionally, the analysis method further includes the following steps: S105. Whenever a processing item is matched, all alternative items other than the processing item are taken as derivative items. According to the processing item, it is determined whether there is a derivative requirement that meets the preset derivative condition. If so, the derivative item corresponding to the derivative requirement is determined, and according to the processing template corresponding to the derivative item, a derivative item processing guide is generated and fed back to the user.
[0038] Among them, "determining whether there is a derivative requirement that meets the preset derivative condition. If so, the derivative item corresponding to the derivative requirement is determined" in S105 specifically includes: By means of a pre-constructed item relationship graph, it is determined whether there is a relationship chain that includes the processing item, and the relationship chain also includes other items other than the processing item, and there is an association relationship between adjacent items in the belonging relationship chain, and there is an association requirement for describing the association relationship; among them, the item relationship graph is used to represent the association relationship between different items. Based on all the documents included in the processing instruction, the association strength between each item in the association chain and the processing item is calculated respectively, and the item with the association strength meeting the preset derivative condition is taken as the derivative item, and the corresponding association requirement is taken as the derivative requirement; among them, the association strength at least includes: the processing completion rate of processing the belonging item by using all the documents included in the processing instruction.
[0039] In implementation, as can be seen from the foregoing, the processing item matched by this application is the only processing item. Then, for other alternative items other than the processing item, this application defines them as derivative items.
[0040] Moreover, in addition, this application also proposes to further screen other derivative items related to the processing item based on the processing item and the preset item relationship graph, and then take the union of all the derivative items determined and screened in the above two times to obtain the final derivative item set. Then, according to the processing template corresponding to the derivative item, a corresponding derivative item processing guide is generated for each derivative item and fed back to the user. The derivative item processing guide is used to describe the documents required for executing the corresponding derivative item and the specific processing content. The specific method for further screening other derivative items related to the processing item based on the processing item and the preset item relationship graph is as follows: The matter relationship graph can be presented in the form of a topological graph, with matters as nodes, and arrows are used to connect matters with associated relationships. The associated requirements are noted at the line segments (e.g., matter A is a prerequisite for matter B), and the associated requirements are used to describe their associated relationships. When the processing matter is "opening registration of a catering store", other matters associated with this processing matter specifically include "food business license", "fire inspection", "application for a health permit", etc. When the processing matter is "company registration", the matter associated with it is "tax registration". Since the realization of "tax registration" itself is a prerequisite for another matter "invoice application", the relationship chain of "company registration → tax registration → invoice application" is correspondingly generated.
[0041] After determining all matters and relationship chains related to the processing matter, the analysis system is used to analyze the association strength between each matter and the processing matter. The calculation rule of the association strength can be predefined manually. Hereinafter, an exemplary calculation method for calculating the association strength will be elaborated: There are several predefined association types and the corresponding association strength values in the analysis system. The association types include: legal mandatory association type (i.e., the government stipulates that certain realizations must be handled in combination, such as "opening registration of a catering store" and "food business license", "fire inspection", "health permit" are legally associated matters); business logic dependency type (i.e., non-legal requirements and there are prerequisite or dependency relationships between matters, such as there is a business dependency relationship between "enterprise loan application" and "mortgage registration", "credit inquiry"); user potential demand type (through historical data mining, matters that most users will derive after handling the processing matter, such as after handling "individual business license", users further handle "tax registration", "social security account opening" matters).
[0042] The analysis system will analyze the association type to which the association relationship between each matter and the processing matter belongs, determine the association strength value, and then compare the document materials required for handling each matter with all the document materials included in the current processing instruction to obtain a similarity value, which is used to represent the processing completion rate. Finally, based on the preset weight value, the association strength value and the similarity value are weighted and summed to obtain a total score. If the total score is higher than the preset score, it is considered that the preset derivative condition is met, and the corresponding matter is taken as the derivative matter of the processing matter.
[0043] Optionally, the analysis method further includes the following steps: Determine whether there are conflict problems between the processing matter and all its corresponding derivative matters. If so, based on the conflict type corresponding to the conflict problem, determine the conflict resolution scheme corresponding to the conflict type, and add the conflict problem and its corresponding conflict resolution scheme to the derivative matter processing guide; Among them, the conflict problem refers to the situation where contradictions occur when the entities shared by the processing matter and the derivative matter are processed according to the corresponding processing content; The conflict types at least include: when the processing matter and the derivative matter execute the corresponding processing content on the shared entity, contradictions occur due to the inconsistent specific content of the shared entity.
[0044] In implementation, the conflict types include conflicts at the data level, that is, when the processing matter and the derivative matter execute the corresponding processing content on the shared entity, contradictions occur due to the inconsistent specific content of the shared entity, that is: the specific values of the same entity field are inconsistent in different files. For example, the user submits: File A "Business License" (including address: No. 1 XX Road), File B "New Lease Contract" (including address: No. 2 XX Road), and based on the foregoing scheme, the processing matter "annual inspection of business license" has been matched for the current case. Then, the alternative implementation "address change record" matched from File B becomes a derivative matter. At this time, there is a conflict in the specific values of the same entity in different files, and the foregoing two files correspond to different matters, thus resulting in a conflict problem between the processing matter and the derivative matter. When executing the processing matter (i.e., processing the annual inspection of the business license), the original address (i.e., No. 1 XX Road) needs to be locked, while when processing the derivative matter (i.e., processing the address change record), the new address (i.e., No. 2 XX Road) needs to be used to overwrite the original address.
[0045] The conflict types can also include conflicts at the rule level, that is, business rules prohibit certain implementation combinations (such as "cancellation" and "social insurance"), so there is a conflict problem between the processing matter "company cancellation" and the derivative matter "social insurance payment".
[0046] The conflict types can also include conflicts at the status level, that is, both matter A and matter B need to use resource X (such as file X), but their required statuses for resource X are different. Exemplarily, the user submits File C "Loan Application Form", and the corresponding processing matter "business loan" is matched, and File D "Contract Record Application", and the corresponding derivative matter "contract record" is matched; when handling the "business loan", the real estate certificate (i.e., resource X) needs to be mortgaged, but when handling the "contract record", the real estate certificate is also needed. That is, when handling the "contract record", it is required that the real estate certificate is in the "idle" state, but when handling the "business loan", the real estate certificate is in the "occupied" state. Therefore, there is a conflict problem between the foregoing "business loan" and "contract record".
[0047] The analysis system can be pre-defined with conflict types and the corresponding conflict cases for each conflict type. The conflict cases include the specific matters in conflict, so as to facilitate the analysis system to analyze whether there are conflict problems after each derived matter is obtained. In addition, for each conflict type, the analysis system also pre-stores the corresponding conflict resolution scheme. Exemplarily, the conflict resolution scheme can include: Priority override strategy: Statutory mandatory matters take precedence over user-initiated matters, which in turn take precedence over the derived matters recommended by the system. For example, address change (statutory matter) is processed prior to business license annual inspection (a matter proposed by the user); it can also include: Interacting with the user to allow the user to select and feedback the solution method by themselves, and implementing conflict resolution based on the solution method feedback by the user.
[0048] Optionally, the analysis method further includes: Regularly, based on the conflict problems discovered in history, the matters related to the conflict problems are used as target matters; Disassemble and reorganize all the processing contents included in the processing template corresponding to the target matter, and satisfy: there are conflict contents in the reorganized processing contents, and the conflict contents are the processing contents that only contain the corresponding conflict problems, and establish and store the conflict relationships among the conflict contents, conflict problems, and the corresponding matters; The conflict resolution scheme at least includes giving priority to processing the processing contents other than the conflict contents.
[0049] In implementation, regularly according to the conflict problems generated in the historical period and the frequency of discovery of the conflict problems, when the frequency is higher than the preset frequency, the corresponding matters are used as target matters, and the processing contents included in the target matters are disassembled and reorganized, that is, each processing content is finally disassembled into processing steps that cannot be further split. The processing steps are the smallest business processing rules that are independent and executable without relying on other processing steps, and the conflict contents are marked out from the disassembled processing contents. The other processing contents that are not conflict contents can be reorganized. Subsequently, whenever a conflict problem related to this conflict content is encountered, the non-conflict contents included in the matter can be processed preferentially, so as to classify the processing contents by disassembling, which is convenient for finding the processing contents that can be processed in parallel and helps to promote the processing progress of the corresponding matters.
[0050] Optionally, the analysis method further includes the following steps: Whenever a processing matter is matched, determine and store the matching relationship between the processing matter and its corresponding key file, where the key file refers to the file that plays a decisive role in the process of matching and determining the processing matter; Regularly analyze the sensitivity of each key document based on the matching relationships stored in historical periods, and mark the key documents with a sensitivity higher than the preset sensitivity threshold as highly sensitive documents. Determine a set of matters for each highly sensitive document, where the set of matters includes the matters that have a matching relationship with the corresponding highly sensitive document. When matching processing matters for a processing instruction, if there is a highly sensitive document among all the documents included in the processing instruction, preferentially retrieve the matters in the set of matters corresponding to the highly sensitive document to perform the matching operation.
[0051] In implementation, combining the specific matching process of the above-mentioned matching processing matters, it can be seen that this application is used to compare the similarity between the file parsing result and the specific file content corresponding to the matter stored in the knowledge base. Then, in this process, the analysis system will establish a "matter-key document" matching relationship. It can be considered that the parsing result extracted from the key document is consistent with the specific file content stored in the corresponding matter. That is to say, all the key documents corresponding to the matter jointly match to obtain the corresponding matter.
[0052] After determining the above matching relationship, the analysis system will determine the number of occurrences of each key document in all matching relationships as the sensitivity of the corresponding key document, and consider that the fewer the number of occurrences, the higher the corresponding sensitivity. When the sensitivity is higher than the preset sensitivity threshold, the corresponding key document is marked as a highly sensitive document. Correspondingly, each time when it is necessary to match processing matters for a processing instruction subsequently, first determine whether there is a highly sensitive document in the processing instruction. If there is, preferentially retrieve the matters that have a matching relationship with the highly sensitive document to perform the matching operation, thereby defining the matching priority order for all matters.
[0053] The embodiment of this application also discloses an adaptive knowledge graph construction and in-depth analysis system that supports multiple types of documents. Refer to Figure 2 , including: A file parsing module 201, configured to receive a processing instruction proposed by a user, and parse all the documents included in the processing instruction. The parsing operation at least includes extracting file entities; A matter matching module 202, configured to match processing matters and corresponding processing templates for the processing instruction based on the parsing result; wherein, the processing template includes processing content for processing entities in the document. A cross-document association module 203, configured to determine and establish an association relationship among all the documents included in the processing instruction based on a preset entity alignment technology. A matter processing module 204, configured to execute all the processing content through a preset processing engine according to the determined processing template and the inter-document association relationship, and feedback the processing result to the user.
[0054] Optionally, it further includes a derivative requirement mining module, which is used to determine whether there is a derivative requirement that meets the preset derivative conditions according to the processing matter every time a processing matter is matched. If so, it determines the derivative matter corresponding to the derivative requirement, and generates and feeds back a derivative matter processing guide to the user according to the processing template corresponding to the derivative matter.
[0055] Optionally, the derivative requirement mining module is further used to determine whether there is a relationship chain containing the processing matter through a pre-constructed matter relationship graph. The relationship chain also contains other matters except the processing matter, and there is an association relationship between adjacent matters in the belonging relationship chain, and there is an association requirement for describing the association relationship; among them, the matter relationship graph is used to represent the association relationship between different matters; it is also used to calculate the association strength between each matter in the association chain and the processing matter respectively based on all the files included in the processing instruction, and use the matter with the association strength meeting the preset derivative conditions as the derivative matter, and use the corresponding association requirement as the derivative requirement; among them, the association strength at least includes: the processing completion rate of using all the files included in the processing instruction to process the belonging matter.
[0056] Optionally, the matter matching module 202 is further used to compare the similarity between the keywords in the parsing result and the data set based on the data set stored in the preset structured database corresponding to each matter, determine all alternative matters and the similarity corresponding to each alternative matter; use the alternative matter with the highest similarity as the processing matter and determine the corresponding processing template, and use all alternative matters other than the processing matter as derivative matters.
[0057] Optionally, it further includes a conflict resolution module, which is used to determine whether there is a conflict problem between the processing matter and all its corresponding derivative matters. If so, it determines the conflict resolution plan corresponding to the conflict type based on the conflict type corresponding to the conflict problem, and adds the conflict problem and its corresponding conflict resolution plan to the derivative matter processing guide; among them, the conflict problem refers to the situation where the entities shared by the processing matter and the derivative matter are processed in contradiction when processed according to the corresponding processing content; the conflict type at least includes: when the processing matter and the derivative matter execute the corresponding processing content on the shared entity, they are contradictory due to the inconsistent specific content of the shared entity.
[0058] Optionally, it further includes a processing content classification module, which is used to regularly use the matters related to the conflict problem as target matters based on the historically discovered conflict problems; disassemble and reorganize all the processing contents included in the processing template corresponding to the target matters, and meet the following conditions: there is conflict content in the reorganized processing content, and the conflict content is only the processing content corresponding to the corresponding conflict problem, and establish and store the conflict relationship between the conflict content, the conflict problem, and the corresponding matters; the conflict resolution plan at least includes giving priority to processing the processing content other than the conflict content.
[0059] Optionally, it further includes a sensitivity matching module, which is used to determine and store the matching relationship between the processing item and its corresponding key file whenever a processing item is matched. Herein, the key file refers to the file that plays a decisive role in the process of matching and determining the processing item. Regularly analyze the sensitivity of each key file based on the matching relationships stored in historical periods, and mark the key files with sensitivity higher than the preset sensitive threshold as highly sensitive files. Determine an item set for each highly sensitive file, and the item set contains the items that have a matching relationship with the corresponding highly sensitive file. When matching processing items for a processing instruction, if there is a highly sensitive file among all the files included in the processing instruction, preferentially retrieve the items in the item set corresponding to the highly sensitive file to perform the matching operation.
[0060] An embodiment of the present application also discloses an adaptive knowledge graph construction and in-depth analysis device supporting multiple types of files. The adaptive knowledge graph construction and in-depth analysis device supporting multiple types of files includes a memory and a processor. A computer program capable of being loaded and executed by the processor, such as the above-mentioned adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of files, is stored on the memory.
[0061] An embodiment of the present application also discloses a computer-readable storage medium, which stores a computer program capable of being loaded and executed by the processor, such as the above-mentioned adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of files. The computer-readable storage medium includes, for example: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0062] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0063] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit the protection scope of the application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all embodiments. Based on these embodiments, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope to be protected by the present application.
Claims
1. An adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of files, characterized in that, Including: Receiving a processing instruction proposed by a user, and parsing all files included in the processing instruction, where the parsing operation at least includes extracting file entities; Based on the parsing result, matching a processing matter and a processing template corresponding to the processing matter for the processing instruction; wherein, the processing content for processing entities in the file is included in the processing template; Based on a preset entity alignment technology, determining and establishing an association relationship among all files included in the processing instruction; According to the determined processing template and the association relationship among files, executing all processing content through a preset processing engine, and feeding back a processing result to the user.
2. The adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of files according to claim 1, characterized in that The method further includes: Whenever a processing matter is matched, determining whether there is a derivative requirement that meets a preset derivative condition according to the processing matter. If so, determining the derivative matter corresponding to the derivative requirement, and generating and feeding back a derivative matter processing guide to the user according to the processing template corresponding to the derivative matter.
3. The adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of files according to claim 2, wherein The determining whether there is a derivative requirement that meets a preset derivative condition, and if so, determining the derivative matter corresponding to the derivative requirement, includes: Determining, through a pre-constructed matter relationship graph, whether there is a relationship chain including the processing matter, where the relationship chain further includes other matters except the processing matter, and there is an association relationship between adjacent matters in the belonging relationship chain, and there is an association requirement for describing the association relationship; wherein, the matter relationship graph is used to represent the association relationship between different matters; Based on all files included in the processing instruction, calculating the association strength between each matter in the association chain and the processing matter respectively, taking the matter with the association strength meeting the preset derivative condition as the derivative matter, and taking the corresponding association requirement as the derivative requirement; wherein, the association strength at least includes: the processing completion rate of processing the belonging matter by using all files included in the processing instruction.
4. The adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of files according to claim 3, characterized in that, The parsing operation further includes performing semantic recognition on the file and extracting keywords; The matching, based on the parsing result, of a processing matter and a processing template corresponding to the processing implementation for the processing instruction includes: Based on the data sets stored in a preset structured database corresponding to each matter, comparing the similarity between the keywords in the parsing result and the data sets to determine all alternative matters and the similarity corresponding to each alternative matter; Taking the alternative matter with the highest similarity as the processing matter and determining the corresponding processing template, and taking all alternative matters except the processing matter as derivative matters.
5. The adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of documents according to claim 4, wherein The method further includes: Determining whether there is a conflict problem between the processing matter and all its corresponding derivative matters. If so, determining a conflict resolution scheme corresponding to the conflict type based on the conflict type corresponding to the conflict problem, and adding the conflict problem and its corresponding conflict resolution scheme to the derivative matter processing guide; Wherein, the conflict problem refers to a situation where there is a contradiction when the entity shared by the processing matter and the derivative matter is processed according to the corresponding processing content. The conflict types at least include: when the processing matters and the derivative matters execute the corresponding processing contents on the shared entity, they are contradictory due to the inconsistent specific contents of the shared entity.
6. The adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of files according to claim 4, wherein The method further includes: Regularly based on the conflict problems discovered in history, regarding the matters related to the conflict problems as target matters; Disassembling and reorganizing all the processing contents included in the processing template corresponding to the target matters, and satisfying: there are conflict contents in the reorganized processing contents, and the conflict contents are the processing contents only including the corresponding conflict problems, establishing and storing the conflict relationships among the conflict contents, the conflict problems, and the corresponding matters; The conflict resolution scheme at least includes preferentially processing the processing contents other than the conflict contents.
7. The adaptive knowledge graph construction and in-depth analysis method for supporting multiple types of files according to claim 1, characterized in that The method further includes: Whenever a processing matter is matched, determining and storing the matching relationship between the processing matter and its corresponding key file, where the key file refers to the file that plays a decisive role in the process of matching and determining the processing matter; Regularly based on the matching relationships stored in the historical period, analyzing the sensitivity of each key file, and calibrating the key files with sensitivity higher than the preset sensitivity threshold as high-sensitivity files, and determining a matter set for each high-sensitivity file, where the matter set includes the matters having a matching relationship with the corresponding high-sensitivity file; When matching a processing instruction with a processing matter, if there is a high-sensitivity file among all the files included in the processing instruction, preferentially retrieving the matters in the matter set corresponding to the high-sensitivity file to perform the matching operation.
8. An adaptive knowledge graph construction and in-depth analysis system supporting multiple types of files, characterized in that, Including, A file parsing module (201) for receiving a processing instruction proposed by a user and parsing all the files included in the processing instruction, and the parsing operation at least includes extracting file entities; A matter matching module (202) for, based on the parsing result, matching a processing matter and a processing template corresponding to the processing matter for the processing instruction; where the processing template includes processing contents for processing entities in the file; A cross-file association module (203) for determining and establishing an association relationship among all the files included in the processing instruction based on a preset entity alignment technology; A matter processing module (204) for, according to the determined processing template and the file association relationship, executing all the processing contents through a preset processing engine and feeding back a processing result to the user.
9. An adaptive knowledge graph construction and in-depth analysis device supporting multiple types of files, characterized in that Including a memory and a processor, and a computer program capable of being loaded and executed by the processor is stored on the memory, and the computer program is the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program capable of being loaded and executed by the processor is stored, and the computer program is the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Knowledge graph construction method, tool, device and server
CN113157947A
Travel guidance generation method and device, equipment and storage medium
CN115238945A
Medical examination report analysis system and method based on clinical examination medical big data
CN115472256A
Policy text denoising and associated item extraction method and system based on large model
CN118820403A
Code map construction method and device, electronic equipment and computer storage medium
CN119149757A