Data resource exploration cataloging method and system for organization digital resource inventory

CN122654099APending Publication Date: 2026-08-28DIGITAL ZHEJIANG TECH OPERATION CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610926218.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

然而,现有技术在处理尚未入库的非结构化或半结构化业务材料(如项目建议书、验收材料、运维清单等)时,存在人工核查工作量大、处理效率低、确认编目后通过率低的问题

Benefits of technology

[0014] This invention provides a data resource exploration and cataloging method and system for organizational digital resource inventory. First, evidence extraction is performed on the business text to be processed, converting it into field-level evidence fragments. Then, ownership determination is performed on these field-level evidence fragments to identify a target exploration knowledge base from a partitioned exploration knowledge base, and the field-level evidence fragments are written into the target exploration knowledge base. Next, based on multiple field-level evidence fragments stored in the target exploration knowledge base, evidence chain asset cards are generated. Finally, after the evidence chain asset cards undergo state-level deduplication, they are sent to a designated associated terminal. The designated associated terminal then responds to the cataloging operation for the evidence chain asset cards, obtaining the cataloged resource results. This method effectively reduces the workload of manual verification, improves processing efficiency, and increases the pass rate after cataloging confirmation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122654099A_ABST
    Figure CN122654099A_ABST
Patent Text Reader

Abstract

The application provides a data resource exploration cataloging method and system for organization digital resource inventory, and relates to the technical field of resource exploration cataloging, which comprises the following steps: extracting evidence from a to-be-processed business text to convert the to-be-processed business text into a field-level evidence segment; determining the ownership of the field-level evidence segment to determine a target exploration knowledge base from a partitioned exploration knowledge base and write the field-level evidence segment into the target exploration knowledge base; generating an evidence chain asset card based on a plurality of field-level evidence segments stored in the target exploration knowledge base; after the evidence chain asset card is state-layered and deduplicated, sending the evidence chain asset card to a designated associated terminal to obtain a cataloging resource result through the designated associated terminal in response to a cataloging operation on the evidence chain asset card. The application can effectively reduce the workload of manual checking, improve the processing efficiency and the pass rate after confirmation cataloging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of resource exploration and cataloging technology, and in particular to a data resource exploration and cataloging method and system for organizational digital resource inventory. Background Technology

[0002] Organizational-level digital resource inventory differs from traditional database metadata scanning. The objects being inventoried are mostly not yet in standard databases or metadata platforms; a large number of high-value leads are scattered across unstructured or semi-structured business materials such as project proposals, design plans, acceptance documents, maintenance checklists, and cloud resource ledgers. However, existing technologies suffer from significant challenges in processing unstructured or semi-structured business materials (such as project proposals, acceptance documents, and maintenance checklists) that have not yet been entered into the database. These challenges include a large workload for manual verification, low processing efficiency, and a low pass rate after confirmation and cataloging. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide a data resource exploration and cataloging method and system for organizational digital resource inventory, which can effectively reduce the workload of manual verification, improve processing efficiency and the pass rate after confirmation and cataloging.

[0004] In a first aspect, the present invention provides a data resource exploration and cataloging method for organizational digital resource inventory, comprising: Evidence extraction is performed on the business text to be processed, so as to convert the business text to be processed into field-level evidence fragments; The ownership of field-level evidence fragments is determined in order to identify the target exploration knowledge base from the partitioned exploration knowledge base, and the field-level evidence fragments are written into the target exploration knowledge base. Based on multiple field-level evidence fragments stored in the target exploration knowledge base, generate evidence chain asset cards; After the evidence chain asset card undergoes state-level deduplication, it is sent to the designated associated terminal so that the designated associated terminal can respond to the cataloging operation on the evidence chain asset card and obtain the cataloging resource result.

[0005] In one implementation, evidence extraction is performed on the business text to be processed to convert it into field-level evidence fragments, including: Identify the target business type corresponding to the business text to be processed; The template matching score is obtained by matching multiple candidate chapter templates associated with the business text to be processed with the target business type, and the target chapter template to be matched with the business text to be processed is determined based on the target matching score. Using the evidence extraction configuration items contained in the target chapter template, evidence is extracted from the business text to be processed, so as to convert the business text to be processed into field-level evidence fragments, which at least contain field confidence levels.

[0006] In one implementation, the method further includes: Determine the evidence priority corresponding to the business text to be processed. The evidence priority is determined based on the authority level of the document type, the authority level of the chapter type, the document version time, the credibility of the field extraction method, and the degree of department matching of the business text to be processed. The template matching score, evidence priority, extraction rule score determined based on evidence extraction configuration items, and OCR confidence decay factor are weighted and fused to obtain the field confidence score.

[0007] In one implementation, the partitioned knowledge base includes at least a global knowledge base and at least one unit knowledge base. The global knowledge base stores field-level evidence fragments that can be accessed by multiple departments, and the unit knowledge base stores field-level evidence fragments that can be accessed by the same department. The process of determining ownership of the field-level evidence fragments to identify the target knowledge base from the partitioned knowledge base includes: Initially define the scope of the evidentiary role of field-level evidence fragments; If the initial calibration fails, write the field-level evidence fragments into the partition to be assigned in the global exploration library; The ownership of field-level evidence fragments stored in the partition to be assigned is determined to ascertain the scope of the evidence fragments. The unit exploration library characterized by the scope of the evidence's scope is used as the target exploration knowledge base, and the scope of the evidence's scope is written back to the field-level evidence fragments.

[0008] In one implementation, ownership determination is performed on field-level evidence fragments stored within the partition to be owned, in order to determine the evidentiary scope of the field-level evidence fragments, including: The field-level evidence fragments stored in the partition to be assigned are matched with the department attribute information corresponding to multiple departments in the preset organizational structure to determine the department matching score; The scope of evidence application for field-level evidence fragments is determined based on department matching scores.

[0009] In one implementation, after generating an evidence chain asset card based on multiple field-level evidence fragments stored in the target exploration knowledge base, the method further includes: Evidence anchoring verification is performed on the evidence chain asset card. Evidence anchoring verification is used to: determine whether the semantic features of the field carried by the evidence fragment identifier in the field contained in the evidence chain asset card are consistent with the semantic features of the original text fragment contained in the field-level evidence fragment associated with the evidence fragment identifier. In addition, the numerical reasonableness of the asset cards in the evidence chain is verified; In addition, cross-validation of evidence is performed on the evidence chain asset card. Cross-validation is used to: determine whether the number of evidence supports for the key field meets the quantity threshold based on the evidence fragment identifier carried by the key field contained in the evidence chain asset card, and / or whether the confidence level of the field contained in the field-level evidence fragment associated with the evidence fragment identifier meets the confidence level threshold.

[0010] In one implementation, the partitioned knowledge base further includes a historical data area, which is used to store at least cataloged resource results and historically unclaimed evidence chain asset cards. The method also includes: Determine the first similarity between the evidence chain asset card and the evidence chain asset card already stored in the target exploration knowledge base; Determine the second similarity between the evidence chain asset card and the historical unclaimed evidence chain asset cards stored in the historical data area; Determine the third similarity between the asset cards in the evidence chain and the cataloged resource results stored in the historical data area; Based on one or more of the first similarity, second similarity, and third similarity, state-level deduplication is performed on the evidence chain asset cards to obtain the state-level deduplication evidence chain asset cards.

[0011] Secondly, the present invention also provides a data resource exploration and cataloging system for organizational digital resource inventory, comprising: The evidence fragment generation module is used to extract evidence from the business text to be processed, so as to convert the business text to be processed into field-level evidence fragments. The ownership determination module is used to determine the ownership of field-level evidence fragments in order to identify the target exploration knowledge base from the partitioned exploration knowledge base and write the field-level evidence fragments into the target exploration knowledge base. The asset card generation module is used to generate evidence chain asset cards based on multiple field-level evidence fragments stored in the target exploration knowledge base; The multi-level deduplication module is used to send the evidence chain asset card to a designated associated terminal after the evidence chain asset card has undergone state-level deduplication. The designated associated terminal then responds to the cataloging operation on the evidence chain asset card and obtains the cataloging resource result.

[0012] Thirdly, the present invention also provides an electronic device including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement any of the methods provided in the first aspect.

[0013] Fourthly, the present invention also provides a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement any of the methods provided in the first aspect.

[0014] This invention provides a data resource exploration and cataloging method and system for organizational digital resource inventory. First, evidence extraction is performed on the business text to be processed, converting it into field-level evidence fragments. Then, ownership determination is performed on these field-level evidence fragments to identify a target exploration knowledge base from a partitioned exploration knowledge base, and the field-level evidence fragments are written into the target exploration knowledge base. Next, based on multiple field-level evidence fragments stored in the target exploration knowledge base, evidence chain asset cards are generated. Finally, after the evidence chain asset cards undergo state-level deduplication, they are sent to a designated associated terminal. The designated associated terminal then responds to the cataloging operation for the evidence chain asset cards, obtaining the cataloged resource results. This method effectively reduces the workload of manual verification, improves processing efficiency, and increases the pass rate after cataloging confirmation.

[0015] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained through the structures particularly pointed out in the description and the drawings.

[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating a data resource exploration and cataloging method for organizational digital resource inventory provided by an embodiment of the present invention; Figure 2 This invention provides an overall technical flowchart of a data resource exploration and cataloging method for organizational digital resource inventory. Figure 3 This is an overall technical flowchart of another data resource exploration and cataloging method for organizational digital resource inventory provided by an embodiment of the present invention; Figure 4This is a schematic diagram of the structure of a data resource exploration and cataloging system for organizational digital resource inventory, provided by an embodiment of the present invention. Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Currently, related technologies provide solutions such as those in patent applications (CN110704455A, CN113254507B, CN115510116A, CN115510116A) and CN113254507B, CN113254507B, and CN115510116A, CN113254507B ...B, CN1132545B, CN1132545B, CN1132545B, CN1132545B, CN1132545B, CN1132545B, CN1132545B, CN1132545B, CN1132545B, CN1132545B, CN1132545B, CN1132545B, CN11325

[0021] However, organizational-level digital resource inventory differs from traditional database metadata scanning. Most of the objects being inventoried have not yet entered standard databases or metadata platforms; a large number of high-value leads are scattered across unstructured or semi-structured business materials such as project proposals, design plans, acceptance documents, maintenance lists, and cloud resource ledgers. Existing solutions face the following four core technical bottlenecks when handling such complex scenarios: (i) The chain of evidence for tracing the source is broken, and the lack of "field-level" anchoring of the original text leads to high costs for manual secondary verification. Existing technical solutions mainly generate catalogs based on technical metadata, database tables, or knowledge graphs, but cannot provide field-level evidence chain support. Even if the system infers a candidate system, it cannot specify the original document chapters and locations corresponding to each key attribute.

[0022] (ii) Blurred isolation boundaries and a lack of multi-dimensional data space division can easily lead to cross-domain data pollution and duplicate recommendations. In actual organizational structures, digital resource inventory has clear permission and reuse boundaries. For example, a global cloud resource list or a province-wide item list is public data that can be reused by multiple units, while a single department's internal operation and maintenance list or project plan should only be effective within the departmental context.

[0023] (iii) The entity disambiguation dimension is too singular and lacks multi-level adaptive deduplication, which can easily lead to missed merging and incorrect merging. In multi-source business materials, the phenomenon of different systems with the same name or different names for the same system is quite common. For example, the same system may appear alternately as project name, abbreviation, old name or platform name.

[0024] (iv) Insufficient engineering tolerance and feedback loop make it difficult to handle low-quality raw materials. In the implementation of the project, low quality scanned documents, broken tables across pages, duplicate headers and footers, and non-standard departmental abbreviations will generate a lot of noise.

[0025] Therefore, existing technologies suffer from problems such as large workload for manual verification, low processing efficiency, and low pass rate after confirmation and cataloging when processing unstructured or semi-structured business materials (such as project proposals, acceptance materials, and maintenance lists) that have not yet been entered into the database.

[0026] Based on this, the present invention provides a data resource exploration and cataloging method and system for organizational digital resource inventory, which can effectively reduce the workload of manual verification, improve processing efficiency and the pass rate after confirmation and cataloging.

[0027] To facilitate understanding of this embodiment, a data resource exploration and cataloging method for organizational digital resource inventory, as disclosed in this embodiment of the invention, will first be described in detail. (See [link to relevant documentation]). Figure 1 The diagram shows a data resource exploration and cataloging method for organizational digital resource inventory. This method mainly includes the following steps S102 to S108: Step S102: Extract evidence from the business text to be processed to convert it into field-level evidence fragments.

[0028] The business texts to be processed refer to organizational-level business materials that have not yet been included in the metadata management platform and exist in unstructured or semi-structured form, including but not limited to original documents such as project proposals, feasibility study reports, preliminary design documents, acceptance materials, operation and maintenance service lists, cloud resource ledgers, and system operation manuals.

[0029] Field-level evidence fragments refer to the smallest verifiable data unit generated based on a single business attribute (such as "application name", "construction department", "launch time", "deployment environment"). These fragments contain a complete set of fields identified by Chinese characters, including: evidence number (evidence_id), source file number (file_id), source file version (file_version), evidence scope (dept_scope), section type (section_type), page number or worksheet position (page_no), text block type (block_type), original text start and end character positions (text_span), table cell range (table_cell_range), raw text fragment (raw_text), normalized field value (normalized_value), extraction method (extract_method), OCR confidence (ocr_confidence), and field confidence (field_confidence).

[0030] In one implementation, the target business type corresponding to the business text to be processed is first identified. Then, multiple candidate chapter templates associated with the business text to be processed are matched with the target business type to obtain a template matching score, so as to determine the target chapter template. Finally, the evidence extraction configuration items contained in the target chapter template are used to extract evidence to obtain field-level evidence fragments.

[0031] Step S104: Determine the ownership of the field-level evidence fragments to identify the target exploration knowledge base from the partitioned exploration knowledge base, and write the field-level evidence fragments into the target exploration knowledge base.

[0032] The partitioned exploration knowledge base refers to three types of storage areas logically divided according to the scope of data application: a global exploration base (used to store public materials that can be reused by multiple departments, such as a province-wide list of items and a global cloud resource ledger), a unit exploration base (used to store materials that are only valid for a specific department, such as a department's maintenance list and acceptance materials), and a historical data area (used to store cataloged resources, historically unclaimed resources, historically rejected resources, push receipts, and manual confirmation records). The global exploration base includes a pending attribution partition, used to temporarily store evidence fragments marked "pending attribution" in their scope of application (dept_scope). The target exploration knowledge base is the final storage area determined after ownership assessment.

[0033] In one implementation, firstly, after the field-level evidence fragment is generated, an initial evidence scope (dept_scope) is marked for the field-level evidence fragment based on the organization to which the file upload channel belongs, the source department field filled in the file metadata, or the department to which the uploader belongs. If the above information is missing or unclear, the evidence scope (dept_scope) is marked as "to be assigned" and written into the "to be assigned" partition of the global exploration library in the partitioned exploration knowledge base. Subsequently, for all field-level evidence fragments with "to be assigned" dept_scope, their department matching score (DeptScore) for each candidate department is calculated. When a department's DeptScore ≥ 0.80, the evidence fragment is confirmed to belong to that department, and it is migrated from the "to be assigned" partition of the global exploration library to the corresponding unit exploration library, while the dept_scope field in the field-level evidence fragment is updated. When there are two or more departments with DeptScores ≥ 0.80, or the difference between the highest and second-highest scores < 0.10, a department conflict (dept_conflict = ...) is marked. If the value is true, the migration will not proceed and the case will be manually reviewed. Finally, the field-level evidence fragments that complete the ownership determination will be written into the corresponding target exploration knowledge base (unit exploration base or global exploration base).

[0034] Step S106: Generate an evidence chain asset card based on multiple field-level evidence fragments stored in the target exploration knowledge base.

[0035] Evidence Chain Asset Card: This refers to a structured record built on the basis of a single candidate digital resource (such as "a certain government affairs app system" or "a certain population database"). It includes fields such as resource ID (entity_id), resource type (resource_type), normalized resource name (normalized_name), entity summary (entity_summary), field value set (field_values), field-level evidence list (field_evidence_list), source document set (source_documents), source section set (source_sections), recommendation cataloging reason (recommend_reason), overall confidence (confidence), conflict flags (conflict_flags), generation batch (batch_id), and generation time (generated_at). Among them, each field value in the field value set (field_values) must be associated with at least one evidence ID (evidence_id), and each record in the field-level evidence list (field_evidence_list) points to the original location and original text of the corresponding evidence fragment, thus forming a complete evidence chain that can be traced field by field.

[0036] In one implementation, all field-level evidence fragments within the same departmental context and that can be referenced in the target exploration knowledge base are taken as input and aggregated according to resource type (e.g., "application system", "data directory"). For each aggregated evidence, rule extraction, named entity recognition, table header mapping, or large language model structured extraction methods are used to generate candidate resource entities. When the model outputs, it is mandatory to return in JSON Schema format, and each field value must be accompanied by one or more evidence numbers (evidence_id). The system verifies whether each field value can find semantically consistent original text evidence in the corresponding evidence number (evidence_id) raw text fragment (using edit distance ≤ 30% of field length or semantic similarity). (≥0.82 judgment); fields that fail the validation are marked as "hallucination_suspected" and written into discarded fields; for fields that pass the validation, the principal value is selected and filled into the field value set (field_values) according to evidence priority (EvidPriority), document version time, field confidence and departmental consistency, and the remaining candidate values ​​are stored in alternative field values ​​(alternative_values); finally, all fields that pass the validation, their associated evidence number (evidence_id) list, source document and chapter information, recommendation reasons and comprehensive confidence, etc., are packaged into a complete evidence chain asset card.

[0037] Step S108: After the evidence chain asset card is deduplicated by state layering, the evidence chain asset card is sent to the designated associated terminal so that the designated associated terminal can respond to the cataloging operation for the evidence chain asset card and obtain the cataloging resource result.

[0038] To facilitate understanding, this invention provides an implementation method for a data resource exploration and cataloging method for organizational digital resource inventory. This method extracts clues of digital resources such as application systems and data directories from multi-source business documents and generates evidence chain asset cards. The method takes multi-source business documents as the core input and outputs evidence chain asset cards with high credibility, traceability, and compliance with publication standards through pipeline-style exploration and extraction, context alignment, and deduplication fusion techniques.

[0039] For details, see Figure 2 The diagram shows the overall technical flow of a data resource exploration and cataloging method for organizational digital resource inventory, and... Figure 3The diagram shows an overall technical flowchart for another data resource exploration and cataloging method for organizational digital resource inventory, including: S1 File access and version indexing; S2 Business file type identification; S3 Chapter template matching; S4 Evidence fragment extraction; S5 Writing evidence fragments to a partitioned exploration knowledge base (including initial labeling of dept_scope); S6 Constructing departmental context and writing back to dept_scope; S7 Generating candidate resource entities; S8 Generating evidence chain asset cards; S9 Multi-level deduplication and fusion; S10 Manual confirmation (including AI-assisted adjustment); S11 Push and receipt synchronization; S12 Boundary sample collection and parameter maintenance. S1-S4 correspond to the exploration processing layer, S5-S9 correspond to the entity generation and fusion layer, and S10-S12 correspond to the confirmation and cataloging closed-loop layer.

[0040] The embodiments of the present invention provide implementation methods for each step.

[0041] The purpose of the aforementioned S1 file access and version indexing is to construct a panoramic knowledge boundary. This stage is responsible for aggregating the static inputs and dynamic reference bases required by the system, including: Multi-source business texts to be processed: As the core dynamic input of the system, these texts cover unstructured or semi-structured documents such as application system design schemes, system operation manuals, bidding documents, and security evaluation reports.

[0042] Organizational Tree and Department Mapping: Provides the authority hierarchy and mapping relationship of organizational structures, serving as the basis for verifying the attribution of exploration results.

[0043] Historical cataloged data: Import existing cataloged resource catalogs as a reference comparison library for subsequent generation of new entities, identification of increments, and conflict compliance determination.

[0044] In this embodiment of the invention, the uploading of business texts to be processed from multiple sources and in multiple formats is supported, and a version control mechanism for files is established to ensure traceability back to the original document version.

[0045] The implementation process for the aforementioned S2 business file type identification (i.e., identifying the target business type corresponding to the business text to be processed) is as follows: In one implementation, the business text to be processed is preprocessed, and the preprocessed result is used as input text. A business file type recognition module automatically identifies the target business type corresponding to the input text. The specific implementation process is as follows: The input to the business document type identification module is the preprocessed document text, table structure, page layout, and document metadata. Supported business document types include cloud resource lists, government affairs lists, system lists, project proposals, feasibility study reports, preliminary designs, detailed designs, acceptance materials, operation and maintenance service lists, and data catalog ledgers.

[0046] The business document type identification module first standardizes the input files: it removes whitespace, full-width / half-width differences, and common numbering symbols from filenames, titles, table headers, chapter titles, and body paragraphs; it normalizes aliases such as "preliminary design," "feasibility study," and "acceptance report / acceptance materials"; and it retains sheet names, first row headers, merged cell ranges, and column name sequences for Excel spreadsheets.

[0047] For each type of business document The system calculates the type matching score. : ; The specific scoring rules for each item are as follows: (a) (Metadata Hit Score): If the file originates from an interface and the type t marked by the interface matches the current candidate type, score 1.0; if the upload channel or manually selected field completely matches type t, score 0.85; if the metadata contains key descriptive words of type t (such as "acceptance", "design", "list", etc.), score 0.5; otherwise, score 0.

[0048] (b) (Title Hit Score): If the file name or homepage title contains the complete standard name of type t (such as "Preliminary Design" or "Feasibility Study Report"), the score is 1.0; if it contains an alias or part of the keywords of type t (such as "Preliminary Design" hitting "Preliminary Design"), the score is 0.6; if only some anchor words appear in the title, the score is linearly interpolated to the range [0, 0.5] based on the ratio of the number of hit anchor words to the total number of anchor words; otherwise, the score is 0.

[0049] (c) (Header Hit Score): If the table header completely contains the set of required headers for type t (such as "System Name", "Cloud Host Name", "Responsible Department"), the score is 1.0; if the required header coverage ratio is greater than or equal to 0.7, the score is linearly interpolated according to the coverage ratio; if the coverage ratio is less than 0.7, the score is 0.5 times the coverage ratio; otherwise, the score is 0.

[0050] (d) (Chapter Hit Score): If the document contains a chapter title that completely matches the standard chapter name of type t (such as "Construction Content", "System Functions", "Acceptance Conclusion"), score 1.0; if the chapter title matches the chapter alias, score 0.7; if only the anchor word is matched but there is no complete chapter title, score 0.35; otherwise, score 0.

[0051] (e) (Anchor Keyword Hit Score): The proportion of anchor keywords corresponding to hit type t in the main text is normalized to the range [0,1]. When an anchor keyword appears in a key position (such as the first paragraph, the first row of a table, or the first sentence of a chapter), the weight of that anchor keyword is multiplied by 1.5.

[0052] (f) (Layout matching score): If the document layout contains the feature layout of type t (such as acceptance materials containing signature pages, feasibility study reports containing table of contents and chapter numbering system, and list-type documents in multi-list format), take 1.0; if only some features match, take 0.5; otherwise, take 0.

[0053] to Type identification weights are assigned, and the sum of all weights is constrained to 1.0. In one embodiment, Take a value between 0.10 and 0.20. Take a value between 0.18 and 0.30. Take values ​​between 0.18 and 0.28. Take a value between 0.15 and 0.25. Take a value between 0.10 and 0.20. Take values ​​from 0.05 to 0.12, while satisfying... The system takes values ​​under the given conditions. The system can adjust the above weights based on historical manually confirmed samples, and the normalization constraint should still be satisfied after adjustment.

[0054] Furthermore, the confidence level of the business document type is determined according to the following priority order: Judgment Rule 1 (High Confidence Hit): When the highest confidence hit... When the score is greater than or equal to 0.78, and the difference between the highest score and the second highest score is greater than or equal to 0.12, the system determines it as a high-confidence hit, directly determines the file_type, and loads the corresponding chapter template.

[0055] Rule 2 (Unrecognizable): When the highest When the difference is less than 0.55, regardless of the size of the difference, the system will mark the file as... Then proceed to the manual specification or general template extraction process.

[0056] Judgment Rule 3 (Medium Confidence Multiple Candidates): When neither Rule 1 nor Rule 2 is satisfied (i.e., the highest confidence multiple candidate) Located within the left-closed, right-open interval of 0.55 to 0.78; or the highest (For scores greater than or equal to 0.78 but with a difference less than 0.12), the system retains the two candidates with the highest scores. Then proceed to multi-template matching.

[0057] The three rules above are mutually exclusive and exhaustive, and are evaluated in the order of rule one, rule two, and rule three.

[0058] The processing after a high-confidence match includes: writing the file type (file_type), type recognition confidence (type_confidence), and type determination basis (type_reason); loading the chapter template set corresponding to the file type (file_type); writing low-priority candidate types into candidate file types (candidate_file_types) for subsequent review; if a high field missing rate or a template matching score is lower than the threshold occurs in the subsequent evidence extraction stage, the system will fall back to multi-template matching.

[0059] Optionally, business file type identification can be implemented using rule scoring, machine learning classification models (including large model classification prompts), or multi-model voting. All implementations should output file_type, type_confidence, and type_reason, and apply different processing strategies for high-confidence hits, low-confidence hits, and unknown types (see other sections of this specification for specific strategies).

[0060] The aforementioned S3 chapter template matching involves matching the business text to be processed with multiple candidate chapter templates associated with the target business type to obtain a template matching score. The target chapter template for matching the business text to be processed is then determined based on the target matching score. The specific implementation process is as follows: Chapter templates are used to describe the chapters, headers, anchor words, and field structures from which evidence can be extracted in a certain type of business document. As shown in Table 1, a chapter template contains the following configuration items: Table 1. Configuration Items for Chapter Templates

[0061] The system calculates the template matching score for each candidate chapter template. : ; This indicates the score for hitting the chapter title or alias; Indicates the coverage of the table header; This indicates the score for anchor word hits; Indicates the score based on layout position, table format, or table of contents level; This indicates the score based on the order of adjacent chapters.

[0062] In one embodiment to Match weights to the template. A value of 0.28 to 0.34 is acceptable. A value of 0.18 to 0.27 is acceptable. A value of 0.16 to 0.23 is acceptable. A value of 0.10 to 0.18 is acceptable. The values ​​can be between 0.05 and 0.12, and the sum of all weights is constrained to 1.0.

[0063] when When the value is greater than or equal to 0.72, the system determines that the chapter has a high confidence match with the corresponding template, binds the chapter template number (template_id), and executes the field extraction rules under that template; when When the value is between 0.55 and 0.72, the system retains multiple candidate templates and performs secondary verification based on field completeness and evidence priority after field extraction; when... When the confidence level is less than 0.55, the system uses a general evidence extraction template and reduces the confidence level of the corresponding field.

[0064] Optionally, chapter template matching can employ rule-based templates, heading hierarchy analysis, layout recognition, semantic retrieval, vector recall, or chapter boundary detection models. Regardless of the implementation method, maintainable template_id, section_type, and field extraction rules should be generated, and evidence_id should be output that can be referenced by evidence chain asset cards.

[0065] The aforementioned S4 evidence fragment extraction involves using the evidence extraction configuration items contained in the target chapter template to extract evidence from the business text to be processed, thereby converting the business text to be processed into field-level evidence fragments. Each field-level evidence fragment contains at least the field confidence level.

[0066] Evidence fragments are the basic data units of this application. The system generates evidence fragments through OCR, text parsing, table parsing, chapter template matching, and field extraction rules.

[0067] Table 2. Core fields and meanings of each piece of evidence.

[0068] Furthermore, embodiments of the present invention also require determining the field confidence level of the evidence fragment, the process of which is as follows: (1) Determine the evidence priority corresponding to the business text to be processed. The evidence priority is determined based on the authority level of the document type, the authority level of the chapter type, the document version time, the credibility of the field extraction method, and the degree of department matching of the business text to be processed. Specifically, the evidence priority... Calculated by weighted scores from the following five factors: ; Table 3 Factor meanings and scoring rules

[0069] It should be noted that the priority of evidence Although the extraction method factor is already included (Weight 0.18), but the score of this method focuses on measuring the overall credibility level of the evidence fragment in the chain of evidence (comprehensively considering five dimensions: document type, chapter, version, extraction method, and department), while the score of the extraction method is less important. This focuses on the independent contribution of differences in the precision of the extraction method itself to the quality of field values. Both measure credibility at different granularities: the former is an evidence-level macro-level assessment, while the latter is a field-level micro-level assessment; therefore, they can reasonably coexist in field confidence levels. The calculation does not involve double counting. exist The weights in the data are decayed by a factor of 0.18, which affects the confidence level of the field. The overall impact is approximately 0.25 × 0.18 = 0.045, and the sampling method score is... The effect is approximately 0.30, and the difference in magnitude between the two indicates that they focus on different dimensions.

[0070] In one embodiment, to As weight, A value of 0.20 to 0.30 is acceptable. A value of 0.20 to 0.30 is acceptable. A value of 0.15 to 0.25 is acceptable. A value of 0.12 to 0.22 is acceptable. The values ​​can be between 0.08 and 0.15, with the sum of all weights constrained to 1.0. In a further preferred embodiment, .

[0071] For example, for the "Construction Department" field, the acceptance report signature page ( Evidence of this type takes precedence over generalized descriptions in the main body of the project proposal. For the "Deployment Environment" field, the list of operation and maintenance services and the list of cloud resources ( ) has higher priority than the planning description in the project proposal. ).

[0072] (2) The template matching score and evidence priority, as well as the extraction rule score determined based on the evidence extraction configuration item and the OCR confidence decay factor are weighted and fused to obtain the field confidence. The specific calculation formula is as follows: ; in, The template matching score M for the chapter containing this evidence fragment; Scoring by extraction method (Manual correction = 1.0, Header mapping = 0.92, Rule extraction = 0.82, Model extraction = 0.68, OCR = 0.55). Prioritize the evidence for this particular piece of evidence; The value is 1.0 for non-OCR extraction and 1.0 for OCR extraction. The original value.

[0073] After evidence fragments are generated, they are written to the evidence table without being directly overwritten. If the same file version is parsed again, the system generates a new parsing batch number (parse_batch_id) and retains the results of the old batch. The parsing batch number (parse_batch_id) is a unique identifier assigned by the system to each complete evidence extraction process performed on the same file (with the same file_id and file_version). It is used to distinguish multiple rounds of parsing results of the same file generated at different times and with different configurations (such as template updates, OCR engine upgrades, and rule adjustments).

[0074] Scenarios that trigger re-parsing within the same version include: after a chapter template is updated, evidence extraction needs to be re-executed on historical files using the new template; after an OCR engine upgrade, historical low-confidence text blocks need to be re-identified; changes in field mapping rules caused by boundary sample recycling require verification of the effect of the changes; and manually initiated forced re-parsing operations. After re-parsing is completed, the system performs a field-level difference comparison between the old and new batches of results and outputs a difference comparison report (diff_report). Manual confirmation is required to determine whether the new batch of results should replace the old batch as input for subsequent entity generation.

[0075] If a field value is manually corrected, a new correction record (correction_record) is created, recording the person who made the correction, the correction time, the original value, the corrected value, and the reason code. This prevents the loss of traceability of evidence for subsequent asset cards.

[0076] The rules for handling low-quality materials are as follows: text blocks with an OCR confidence score below 0.75 are marked as low-confidence evidence; for tables spanning multiple pages, whether to merge them is determined based on duplicate headers, column consistency, continuity of the first column number, and row semantic closure; headers and footers are filtered by duplicate text hashes and page position ranges; merged cells are filled with field values ​​by filling in the cell range.

[0077] The aforementioned S5 evidence fragments are written into the partitioned exploration knowledge base (including initial annotations in dept_scope) and the S6 construction department context is written back to dept_scope.

[0078] The partitioned knowledge base includes a global knowledge base, at least one unit knowledge base, and a historical data area. The global knowledge base stores field-level evidence fragments that can be accessed by multiple departments, such as a global cloud resource list, a unified item list, and a unified system list. The unit knowledge base stores field-level evidence fragments that can be accessed by the same department, such as a departmental maintenance list and departmental acceptance materials. The historical data area stores at least cataloged resource results, historically unclaimed evidence chain asset cards, historically rejected resources, push receipts, and manual confirmation records. Optionally, the partitioned knowledge base can use physical isolation or logical isolation achieved through tenant identifiers, department identifiers, and data domain identifiers. The underlying layer can be a relational database, document database, search engine, vector database, or graph database. The above isolation methods and storage selections do not affect the core processing logic of "reusable global materials, restricted departmental materials, and historical status participating in deduplication." Those skilled in the art can choose the appropriate storage method based on the data scale, query performance, and permission isolation requirements of the business system.

[0079] The global exploration repository includes a pending-attribution partition for temporarily storing evidence fragments marked as pending (pending_dept) within the S5 stage's evidence scope (dept_scope). Once the S6 departmental context is constructed, these evidence fragments will be migrated to the corresponding unit exploration repository or confirmed to remain in the global exploration repository.

[0080] Specifically, the scope of the evidentiary effect of field-level evidence fragments is initially defined. If the initial definition fails, the field-level evidence fragments are written into the unassigned partition of the global exploration library. Ownership determination is performed on the field-level evidence fragments stored in the unassigned partition to determine their evidentiary scope. The unit exploration library representing the departmental scope of the evidentiary effect is used as the target exploration knowledge base, and the departmental scope of the evidentiary effect is written back to the field-level evidence fragments. The ownership determination process involves matching the field-level evidence fragments stored in the unassigned partition with the departmental attribute information corresponding to multiple departments in a preset organizational structure to determine the departmental matching score. Based on the departmental matching score, the evidentiary scope of the field-level evidence fragments is determined.

[0081] In this embodiment of the invention, the two-stage department labeling mechanism of S5 and S6 is as follows: S5 writes evidence fragments into the partitioned exploration knowledge base: When writing evidence fragments, dept_scope adopts an initial attribution labeling strategy—initial labeling is based on the organization to which the file upload channel belongs, the source department field in the file metadata, or the department affiliation of the uploader. If the above information is missing or unclear, dept_scope is temporarily marked as pending_dept and written to the pending attribution partition of the global exploration database.

[0082] S6 constructs department context and writes back dept_scope: After the evidence fragment is written, the system calculates DeptScore using the organizational tree and department mapping, the department alias table, and department clues in the document body. It then performs attribution determination on all evidence fragments with dept_scope set to pending_dept. If the determination passes, the evidence fragment is migrated from the pending partition in the global exploration library to the corresponding unit exploration library, and the dept_scope field is updated.

[0083] The two-stage design described above ensures that S5 can complete the write without relying on the complete department context, while S6 can perform the globally optimal department affiliation determination under the condition that all evidence is visible.

[0084] Furthermore, the process of determining the department matching score is as follows: When the system constructs the department context, it calculates the department matching score for each candidate department d. The calculation formula is as follows: ; Table 4 Meaning of each item

[0085] In one embodiment, to As weight, Take a value between 0.35 and 0.45. Take a value of 0.25 to 0.30. Take a value between 0.10 and 0.15. Take a value between 0.08 and 0.12. The weights are set to 0.05 to 0.10, and the sum of all weights is constrained to 1.0.

[0086] when When the value is greater than or equal to 0.80, the system determines that the candidate resource belongs to department d as a high-confidence hit; when When the value is between 0.55 and 0.80, it is retained as a candidate for attribution and the departmental attribution is marked as uncertain. ;when If the value is less than 0.55, the candidate is excluded.

[0087] If the same candidate resource is available to two or more departments Both are greater than or equal to 0.80, or the scores of two candidate departments are... If the difference is less than 0.10, there is a conflict in the system settings regarding department affiliation. (and prohibit subsequent automatic merging and automatic push),

[0088] For the aforementioned S7, candidate resource entities are generated.

[0089] The resource entity generation module generates candidate resource entities based on the current department context, the set of referenceable evidence_ids, and the resource type definition. Generation methods can include rule extraction, named entity recognition, header mapping, semantic modeling, or structured extraction from a large model.

[0090] When using a large model, the input prompts include the allowed evidence_id, field output structure, resource type definition, rules against fabrication, and conflict handling rules. The output must conform to JSONSchema.

[0091] The system performs three types of checks on the output results: field structure check, evidence number existence check, and department context check. Fields that fail the checks are moved to discarded fields or secondary check fields and are not moved to the formal asset card.

[0092] The aforementioned S8 generates a chain of evidence asset cards.

[0093] This invention further performs evidence anchoring verification and no-evidence field blocking on the generated evidence chain asset card, including: Evidence anchoring verification is performed on the evidence chain asset cards. This verification determines whether the semantic features of a field, based on the evidence fragment identifier carried by its fields, are consistent with the semantic features of the original text fragment contained within the field-level evidence fragment associated with that identifier. For example, each formally written field value output by the model should be accompanied by one or more `evidence_id` values. The system verifies whether the output field has original text evidence based on the original text fragment (raw_text) corresponding to the `evidence_id`, the normalized field value (normalized_value), and the field type. The system verifies whether there is an original text fragment in the `raw_text` corresponding to the `evidence_id` that is semantically consistent with the output field value. Consistency determination uses edit distance (character-level) or semantic similarity (vector cosine), with thresholds of edit distance ≤ 30% of the field value length or semantic similarity ≥ 0.82, respectively. Fields that do not meet the conditions are marked as "hallucination_suspected," indicating that the field value lacks verifiable original text evidence in the original document and is speculative content generated by the model or rules without sufficient support, posing a risk of factual error.

[0094] Perform numerical validity checks on the asset cards in the evidence chain. For example, verify the format and time range of date fields (e.g., the online time is not earlier than the project initiation time), and verify the consistency of units and the reasonableness of magnitude of numerical fields.

[0095] Cross-validation of evidence is performed on the evidence chain asset cards. This cross-validation is used to: determine whether the number of pieces of evidence supporting a key field meets a threshold based on the evidence fragment identifier carried by the key field in the evidence chain asset card; and / or whether the field confidence level of the field contained in the field-level evidence fragment associated with the evidence fragment identifier meets a confidence level threshold. For example, when cross-validating a key field (such as system name or construction department) among multiple pieces of evidence, if only a single piece of evidence supports it and the field confidence level (field_confidence) of that evidence is less than 0.70, the confidence level of that field is reduced and marked as weakly supported (weak_evidence_only). This indicates that the field value is supported only by a single-source field-level evidence fragment, and the field confidence level (field_confidence) of that evidence fragment is less than 0.70.

[0096] Fields marked with "hallucination_suspected" are not included in the formal field_values, but are written to the discarded field set, with the reason for discard being recorded as "evidence anchoring verification failed", i.e., discard_reason="hallucination_check_failed".

[0097] Table 5. Evidence Chain Asset Card includes the following main fields.

[0098] Optionally, the field structure of the evidence chain asset card can be adjusted according to the implementation scenario. For example, Word or PDF documents can use page numbers, paragraph numbers, and text_span to identify the source location; Excel ledgers can use sheet_name, row_index, and column_name to identify the source location; field evidence can also be split into a source document evidence table and an inference evidence table. Candidate resource fields can be associated with evidence_id and can be traced back to the source file version and source document location. As long as the field structure can achieve the function of tracing back to a fixed position (such as character level, cell level, or paragraph level) from the current evidence card, regardless of the specific location identifier used, it is an equivalent implementation of the evidence chain card structure described in this application.

[0099] Overall confidence level of the evidence chain asset cards The calculation formula is as follows: ; in, Field confidence for all filled fields Mean; Define the ratio of the number of filled fields to the number of required fields in the resource_type; To support the normalized value of the number of supporting evidence source documents (1.0 for 3 or more documents, 0.75 for 2 documents, and 0.50 for 1 document). Department matching score Normalized value.

[0100] Each field is associated with an `evidence_id`. If multiple candidate values ​​exist for the same field, the system selects the primary value based on evidence priority, file version time, field confidence, and departmental consistency, and writes the other candidate values ​​into the alternative field values ​​(`alternative_values`). It also records the conflict reason (`conflict_reason`). For example, if the "Construction Department" field comes from both the cloud resource inventory and the acceptance report, and both point to the same department after normalization, the evidence is merged; if they are inconsistent, a field-level conflict is set (`field_conflict=true`) and the system proceeds to manual review.

[0101] The specific process for the aforementioned S9 multi-level deduplication fusion is as follows: The multi-level deduplication and fusion module does not put all candidate resources into the same duplicate pool, but instead performs three types of comparisons sequentially according to the object state: The first category: Determine the first similarity between the evidence chain asset cards and the evidence chain asset cards already stored in the target exploration knowledge base. That is, compare the evidence chain asset cards within the current exploration results. If the automatic duplication condition is met and there are no hard constraints, then merge the cards and their field-level evidence chains.

[0102] The second category: Determine the second similarity between the evidence chain asset card and the historical unclaimed evidence chain asset cards stored in the historical data area. That is, compare the current investigation results with the historical unclaimed results. If a historical unclaimed object is matched, the historical evidence chain, processing opinions, and reasons for unclaiming are inherited, and the review priority is increased.

[0103] The third category: Determine the third similarity between the evidence chain asset card and the cataloged resource results stored in the historical data area. That is, compare the current exploration results with the cataloged results. If a cataloged object is matched, mark it with suspected_cataloged and catalog_target_id. Do not generate a new push task, but enter the update confirmation or manual review process.

[0104] Then, based on one or more of the first similarity, second similarity, and third similarity, state-level deduplication is performed on the evidence chain asset cards to obtain the state-level deduplication evidence chain asset cards.

[0105] The similarity calculation formula is as follows : ; Wherein, represents the similarity of standard names, abbreviated dictionaries and historical aliases; represents the semantic similarity of entity abstracts; S_matter represents the similarity of related matter codes, matter names and matter levels; represents the coincidence degree of project numbers, project names and construction periods; represents the matching degree of unified social credit codes, department codes and department names; represents the coincidence degree of source file types, chapter types, evidence positions and field sources.

[0106] For application entities, a1 can be 0.25 to 0.35, a2 can be 0.16 to 0.24, a3 can be 0.10 to 0.18, a4 can be 0.07 to 0.13, a5 can be 0.12 to 0.20, a6 can be 0.07 to 0.15, and the sum of all weights is constrained to 1.0. A trial operation configuration is a1=0.31, a2=0.19, a3=0.14, a4=0.09, a5=0.17, a6=0.10. For data catalog entities, the weights of abstracts and evidence sources are relatively high, a trial operation configuration is a1=0.21, a2=0.27, a3=0.19, a4=0.06, a5=0.11, a6=0.16, and the sum of all weights is constrained to 1.0.

[0107] The first threshold T1 is used for automatic duplicate determination, and its trial operation configuration is 0.86; the second threshold T2 is used for suspected duplicate determination, and its trial operation configuration is 0.64.

[0108] If S is greater than or equal to T1, and there is no department conflict, key field conflict and historical rejection mark, the system automatically merges the evidence chains; if S is between T2 and T1 (that is, T2≤S<T1), it enters manual review; if S is less than T2, it is retained as a new candidate resource. Department conflict (dept_conflict=true), key field conflict (field_conflict=true involves system name or construction department field) and historical rejection mark (the catalog_target_id has a reject_reason record in the historical data area) are hard constraints. Even if S is greater than or equal to T1, automatic merging is not performed, and forced entry into manual review is required.

[0109] For the aforementioned S10 manual confirmation (including AI-assisted adjustment), S11 push and receipt synchronization, and S12 boundary sample collection and parameter maintenance. The specific process is as follows: AI-assisted adjustment function: In the manual confirmation interface, when reviewers view the evidence chain asset card, the system can automatically generate modification suggestions based on the reviewer's operational context (such as the field currently being reviewed and the marked modification comments). Specifically, the system takes the current card's field_values, field_evidence_list, conflict_flags, and the reason_code entered by the reviewer as input, and calls the large model to generate candidate modification schemes, including but not limited to: suggesting changing the primary value to a certain alternative_value, suggesting adding a certain field value, and suggesting changing the department of ownership.

[0110] This function only generates pending_change records, which include suggested_field, suggested_value, suggested_evidence_id, and suggested_reason. The pending_change record represents a temporary status marker for a field in a field-level evidence fragment or evidence chain asset card, indicating that the field value has been confirmed by a human reviewer to require modification, but the correction operation has not yet been completed or the system has not yet performed an update; it is in an intermediate state of "change request initiated, not yet effective."

[0111] The pending_change is not written to the field_values ​​of the formal asset card unless the "confirm acceptance" operation is clicked manually, ensuring human-in-the-loop decision-making power.

[0112] The manual confirmation module receives asset cards generated by the system and supports confirmation (approval after modification), rejection, and transfer to relevant departments for confirmation. Each operation records the operator's unique identifier (operator_id), operation timestamp (operation_time), operation type code (operation_type), before-operation status hash (before_hash), after-operation status hash (after_hash), operation reason code (reason_code), and audit batch number (audit_batch_id).

[0113] The push module converts the confirmed asset card into the target business system interface format, generating the target business system identifier `target_system_id`, the unique identifier of the push resource `resource_id`, the push batch number `push_batch_id`, the payload hash value `payload_hash`, the retry count counter `retry_count`, and the push status code `push_status`. The receipt synchronization module performs idempotent updates based on `target_system_id`, `resource_id`, and `push_batch_id`. When the business system returns a rejection, the system records the `reject_reason` and writes the reason to the historical data area, increasing the review priority when similar resources reappear.

[0114] Table 6. The boundary sample maintenance module collects samples according to the following triggering conditions:

[0115] The system records for each boundary sample a unique identifier (sample_id), a sample type code (sample_type), a trigger condition value (trigger_condition_value), a set of related evidence IDs (related_evidence_ids), a related candidate entity identifier (related_entity_id), and a collection timestamp (collect_time). After manual review, a review_result is generated (with values ​​of confirmed_correct, confirmed_error, and ambiguous). Samples of the confirmed_error class are added to the model fine-tuning training set, samples of the confirmed_correct class are added to the threshold calibration validation set, and samples of the ambiguous class are added to the expert secondary review queue.

[0116] Applications of boundary samples include: updating chapter templates (extending section_alias and anchor_terms), correcting field mapping rules, supplementing the department alias table, fine-tuning evidence priority weights p1–p5, adjusting deduplication weights a1–a6, and calibrating thresholds T1 and T2.

[0117] Through the above approach, the system can convert resource clues in unstructured or semi-structured business materials into evidence fragments with source location, departmental scope, and field confidence levels, and further generate field-level traceable evidence chain asset cards. Compared with methods that only generate catalog entries or metadata entities, this application can preserve the binding relationship between field values ​​and original evidence during the candidate resource generation stage.

[0118] In a set of anonymized verification data, the system collected and processed 883 business materials from the company's business departments, involving 23 departments, and extracted 8,427 candidate evidence fragments. After departmental context filtering, multi-level deduplication, and manual confirmation, 3,194 candidate resource entities were formed, with a manual first-time confirmation pass rate of 77.6%.

[0119] Based on the foregoing embodiments, this invention provides a data resource exploration and cataloging system for organizational digital resource inventory, see [link to relevant documentation]. Figure 4 The diagram shown illustrates the structure of a data resource exploration and cataloging system for organizational digital resource inventory, including: The evidence fragment generation module 402 is used to extract evidence from the business text to be processed, so as to convert the business text to be processed into field-level evidence fragments. The ownership determination module 404 is used to determine the ownership of field-level evidence fragments in order to identify the target exploration knowledge base from the partitioned exploration knowledge base and write the field-level evidence fragments into the target exploration knowledge base. The asset card generation module 406 is used to generate evidence chain asset cards based on multiple field-level evidence fragments stored in the target exploration knowledge base. The multi-level deduplication module 408 is used to send the evidence chain asset card to a designated associated terminal after the evidence chain asset card has undergone state-level deduplication, so that the designated associated terminal can respond to the cataloging operation on the evidence chain asset card and obtain the cataloging resource result.

[0120] The system provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the system embodiment can be referred to the corresponding content in the aforementioned method embodiment.

[0121] This invention provides an electronic device, specifically, the electronic device includes a processor and a memory; the memory stores a computer program, which, when run by the processor, executes the method described in any of the above embodiments.

[0122] Figure 5 The present invention provides a schematic diagram of the structure of an electronic device 100, which includes a processor 50, a memory 51, a bus 52 and a communication interface 53. The processor 50, the communication interface 53 and the memory 51 are connected through the bus 52. The processor 50 is used to execute executable modules, such as computer programs, stored in the memory 51.

[0123] The memory 51 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 53 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.

[0124] Bus 52 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0125] The memory 51 is used to store programs. After receiving an execution instruction, the processor 50 executes the program. The method executed by the system defined by the flow process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 50 or implemented by the processor 50.

[0126] Processor 50 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 50 or by instructions in software form. Processor 50 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 51. The processor 50 reads the information in memory 51 and, in conjunction with its hardware, completes the steps of the above method.

[0127] The computer program product of the readable storage medium provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, please refer to the foregoing method embodiments, which will not be repeated here.

[0128] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0129] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A data resource exploration and cataloging method for organizational digital resource inventory, characterized in that, include: Evidence extraction is performed on the business text to be processed to convert it into field-level evidence fragments; The ownership of the field-level evidence fragments is determined to identify the target exploration knowledge base from the partitioned exploration knowledge base, and the field-level evidence fragments are written into the target exploration knowledge base. Based on the multiple field-level evidence fragments stored in the target exploration knowledge base, an evidence chain asset card is generated; After the evidence chain asset card undergoes state-level deduplication, the evidence chain asset card is sent to a designated associated terminal so that the designated associated terminal can respond to the cataloging operation for the evidence chain asset card and obtain the cataloging resource result.

2. The data resource exploration and cataloging method for organizational digital resource inventory according to claim 1, characterized in that, Evidence extraction is performed on the business text to be processed to convert it into field-level evidence fragments, including: Identify the target business type corresponding to the business text to be processed; The template matching score is obtained by matching the business text to be processed with multiple candidate chapter templates associated with the target business type, and the target chapter template to be matched with the business text to be processed is determined based on the target matching score. Using the evidence extraction configuration items contained in the target chapter template, evidence is extracted from the business text to be processed, so as to convert the business text to be processed into field-level evidence fragments, wherein the field-level evidence fragments contain at least field confidence.

3. The data resource exploration and cataloging method for organizational digital resource inventory according to claim 2, characterized in that, The method further includes: The evidence priority corresponding to the business text to be processed is determined based on the authority level of the file type, the authority level of the chapter type, the file version time, the credibility of the field extraction method, and the degree of department matching of the business text to be processed. The confidence level of the field is obtained by weighting and fusing the template matching score, the evidence priority, the extraction rule score determined based on the evidence extraction configuration item, and the OCR confidence decay factor.

4. The data resource exploration and cataloging method for organizational digital resource inventory according to claim 1, characterized in that, The partitioned exploration knowledge base includes at least a global exploration library and at least one unit exploration library. The global exploration library is used to store field-level evidence fragments that can be accessed by multiple departments, and the unit exploration library is used to store field-level evidence fragments that can be accessed by the same department. Determine the ownership of the field-level evidence fragments to identify the target exploration knowledge base from the partitioned exploration knowledge base, including: The scope of the evidentiary effect of the field-level evidence fragments is initially defined; If the initial calibration fails, the field-level evidence fragment will be written into the partition to be assigned in the global exploration library; The ownership of the field-level evidence fragments stored in the partition to be assigned is determined to ascertain the scope of the evidence. The unit exploration library characterized by the scope of the evidence scope is used as the target exploration knowledge base, and the scope of the evidence scope is written back to the field-level evidence fragments.

5. The data resource exploration and cataloging method for organizational digital resource inventory according to claim 4, characterized in that, To determine the scope of evidence validity of the field-level evidence fragments stored within the partition to be assigned, the following steps are performed: The field-level evidence fragments stored in the partition to be assigned are matched with the department attribute information corresponding to multiple departments in the preset organizational structure to determine the department matching score; The scope of the evidence for the field-level evidence fragment is determined based on the department matching score.

6. The data resource exploration and cataloging method for organizational digital resource inventory according to claim 1, characterized in that, After generating an evidence chain asset card based on multiple field-level evidence fragments stored in the target exploration knowledge base, the method further includes: The evidence chain asset card is subjected to evidence anchoring verification, which is used to: determine whether the semantic features of the field carried by the evidence fragment identifier in the field contained in the evidence chain asset card are consistent with the semantic features of the original text fragment contained in the field-level evidence fragment associated with the evidence fragment identifier. In addition, the numerical reasonableness of the evidence chain asset cards is verified; In addition, cross-validation of evidence is performed on the evidence chain asset card. The cross-validation of evidence is used to: determine whether the number of evidence supports for the key field meets the quantity threshold based on the evidence fragment identifier carried by the key field contained in the evidence chain asset card, and / or whether the confidence level of the field contained in the field-level evidence fragment associated with the evidence fragment identifier meets the confidence level threshold.

7. The data resource exploration and cataloging method for organizational digital resource inventory according to claim 1, characterized in that, The partitioned knowledge base also includes a historical data area, which is used to store at least cataloged resource results and historically unclaimed evidence chain asset cards. The method further includes: Determine the first similarity between the evidence chain asset card and the evidence chain asset cards already stored in the target exploration knowledge base; Determine the second similarity between the evidence chain asset card and the historical unclaimed evidence chain asset cards stored in the historical data area; Determine the third similarity between the evidence chain asset card and the cataloged resource results stored in the historical data area; Based on one or more of the first similarity, the second similarity, and the third similarity, the evidence chain asset card is subjected to state-level deduplication to obtain the evidence chain asset card after state-level deduplication.

8. A data resource exploration and cataloging system for organizational digital resource inventory, characterized in that, include: The evidence fragment generation module is used to extract evidence from the business text to be processed, so as to convert the business text to be processed into field-level evidence fragments. The ownership determination module is used to determine the ownership of the field-level evidence fragments in order to identify the target exploration knowledge base from the partitioned exploration knowledge base and write the field-level evidence fragments into the target exploration knowledge base. The asset card generation module is used to generate an evidence chain asset card based on multiple field-level evidence fragments stored in the target exploration knowledge base; The multi-level deduplication module is used to send the evidence chain asset card to a designated associated terminal after the evidence chain asset card has undergone state-level deduplication, so that the designated associated terminal can respond to the cataloging operation for the evidence chain asset card and obtain the cataloging resource result.

9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for generating data asset directory, terminal and storage medium

    CN110704455A

  • A method for intelligently constructing and inventorying a data asset catalog

    CN113254507B

  • Data directory construction method and device, medium and equipment

    CN115510116A