Intelligent service file management method and system

By performing optical character recognition and natural language processing on contract documents, a standardized coded dataset is formed, which solves the problem of inconsistent document formats in the management of business files in multiple departments, realizes intelligent file management and early warning mechanism, and improves the efficiency of file retrieval and the timeliness of document management.

CN120179885BActive Publication Date: 2025-11-07GUANGZHOU DAIBEI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510243567.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-11-07
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

In existing technologies, the lack of unified standards for managing business files across multiple departments leads to significant differences in file formats and classifications, making retrieval difficult, search inefficient, and updates untimely, with no ability to provide alerts for file issues.

Method used

Contract documents are encoded using optical character recognition technology to form a standard encoded dataset. Natural language processing technology is used to enable rapid document retrieval and early warning prompts, and a dynamically updated index file is established to generate business file early warning information.

Benefits of technology

It enables unified management and real-time updates of various file types, improves the efficiency of file retrieval, ensures the integrity, accuracy, security and compliance of files, and promptly alerts users to abnormal situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179885B_ABST
    Figure CN120179885B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent service file management method and system, which encodes contract files in service files, configures project files and financial files corresponding to the contract files according to the encoding of the contract files, and standardizes the encoding of the files to form a standard encoding data set, reads the standard encoding data set to obtain file information, generates service file early warning information according to the file information, and realizes the unified management of the service files through the standardization of the encoding of the identification information and the formation of the standard encoding data set.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of archive management, in particular to an intelligent business archive management method and system. BACKGROUND

[0002] In the current business activities, cooperation and exchange among multiple departments are involved, and each department has its own work habits and file management methods. In addition to contract files, there are differences in format and classification among different departments for project files, financial files, etc. Different business departments lack unified standards in file format, classification, numbering, etc., which makes it easy for the archive management department to appear chaotic when sorting and archiving, affecting the retrieval efficiency and utilization effect of archives.

[0003] When it comes to business archive management of multiple types of files, the existing technology often cannot use a unified business file standard for business archive management, making it difficult to find files. When searching for related file content, the file index is incomplete, leading to retrieval difficulties, outdated file search methods, low search efficiency, and the inability to update files in a timely manner and provide file problem reminders. SUMMARY

[0004] The present application provides an intelligent business archive management method and system to solve the above problems in the prior art.

[0005] The present application is as follows:

[0006] An intelligent business archive management method, characterized by the following steps:

[0007] Step one: read the contract files in the business archives, and encode the contract files;

[0008] Step two: configure the project files and financial files corresponding to the contract files through the encoding of the contract files;

[0009] Step three: standardize the encoding of the files to form a standard encoding dataset;

[0010] Step four: read the standard encoding dataset to obtain file information;

[0011] Step five: generate business archive warning information according to the file information.

[0012] The step one: reading the contract files in the business archives and encoding the contract files includes: scanning the contract files through optical character recognition technology, identifying the text information of the contract files, encoding the text information, and forming a first encoding dataset S1.

[0013] The first encoding data set S1 is represented as:

[0014]

[0015] Wherein, ID represents a unique identification, including contract name and number; T1 represents time attribute, including signing date, contract period; P1 represents space attribute, including contract associated project, project location; M1 represents asset attribute, including contract transaction price; D1 represents contract content set, including contract terms.

[0016] The second step: through the encoding of the contract file, configuring the project file and the financial file corresponding to the contract file includes:

[0017] Reading the encoding data of the contract file, finding the project file related to it through the encoding data, and encoding the project file to form a second encoding data set S2;

[0018] Reading the encoding data of the contract file, finding the financial file related to it through the encoding data, and encoding the financial file to form a third encoding data set S3.

[0019] The second encoding data set S2 is represented as:

[0020]

[0021] Wherein, ID represents a unique identification, including contract name and number; T2 represents time attribute, including project start time, project end time; P2 represents space attribute, including project scope; M2 represents asset attribute, including project budget; D2 represents project file content set, including project plan, project execution, project monitoring;

[0022] The third encoding data set S3 is represented as:

[0023]

[0024] Wherein, ID represents a unique identification, including contract name and number; T3 represents time attribute, including actual use time of funds; P3 represents space attribute, including the use location of funds; M3 represents asset attribute, including financial statements; D3 represents financial file content set, including profitability analysis, operational capability analysis, financial decision-making;

[0025] The third step: standardizing the encoding of the file to form a standard encoding data set includes:

[0026] According to the first encoding data set S1, the second encoding data set S2 and the third encoding data set S3, a unique identification ID is set for the contract file and the associated project file and financial file, the time attribute T, the space attribute P, the asset attribute M and the file content set D of the file are obtained through the unique identification ID, and a standard encoding data set S is formed, wherein .

[0027] The standard encoding of the file also includes creating an index file, and dynamically updating the standard encoding data set S:

[0028] A special encoding index file data set is set, in the index file, the ID code, the file name, the file type, the main content summary and the associated file ID code information of each file are listed, when a new file is generated or the association relationship between files changes, the mapping relationship in the index file is updated in time, and is recorded in the standard encoding data set S.

[0029] The fourth step: reading the standard encoding data set to obtain file information includes: in the standard encoding data set S, reading the unique identification ID of the contract file to find the associated file content, automatically displaying other files associated with the unique identification ID, reading the text by using natural language technology, and obtaining the file content according to the demand.

[0030] The fifth step: generating business archive early warning information according to the file information includes:

[0031] According to the obtained file content, the early warning conditions of the file on completeness, accuracy, security, compliance and timeliness are set, and early warning is prompted when an abnormal situation occurs.

[0032] The application also provides an intelligent business archive management system, characterized in that it is used to realize the above-mentioned intelligent business archive management method, and the system comprises a reading module, a configuration module, an encoding module, a processing module and an early warning module.

[0033] The reading module reads the contract file in the business archive, and encodes the contract file; the configuration module configures the project file and the financial file corresponding to the contract file through the encoding of the contract file; the encoding module performs standard encoding on the file to form a standard encoding data set; the processing module reads the standard encoding data set to obtain file information; and the early warning module generates business archive early warning information according to the file information.

[0034] Compared with the prior art, the embodiment of the application has the following beneficial effects:

[0035] The application provides an intelligent business archive management method and system, which realizes the unified management of business archives by standardizing the coding of contract files, project files and financial files involved in the business activity process through identification information, forming a standard coding data set; meanwhile, an index file is created, the standard coding data set is dynamically updated, and the real-time update of business archives is realized on the basis of unified management; the natural language technology is adopted to realize the rapid and accurate search of business archives, the early warning information is generated for the specific problems of the files in the business archives and timely reminders are given, and the intelligent management of business archives of various types of associated files is realized. BRIEF DESCRIPTION OF DRAWINGS

[0036] Figure 1 is a flowchart of an intelligent business archive management method provided by an embodiment of the application;

[0037] Figure 2 is a schematic diagram of an intelligent business archive management system provided by an embodiment of the application. DETAILED DESCRIPTION

[0038] The application will be described in detail below with reference to the drawings.

[0039] Embodiment 1

[0040] The application provides an intelligent business archive management method, and a flowchart of the method is as shown in Figure 1 The method comprises the following steps:

[0041] Step 1: reading contract files in business archives and coding the contract files; step 2: configuring project files and financial files corresponding to the contract files through the coding of the contract files; step 3: standardizing the coding of the files to form a standard coding data set; step 4: reading the standard coding data set to obtain file information; and step 5: generating business archive early warning information according to the file information.

[0042] The step 1: reading contract files in business archives and coding the contract files comprises: scanning the contract files through an optical character recognition technology, identifying the text information of the contract files, coding the text information to form a first coding data set S1.

[0043] Specifically, the text information comprises a contract name, a contract number, a signing date, a contract term, a contract associated project, a project location, a contract transaction price and a contract clause.

[0044] The first coding data set S1 is expressed as:

[0045]

[0046] Wherein, ID represents the unique identification, including contract name and number; T1 represents time attribute, including signing date, contract period; P1 represents space attribute, including contract associated project, project location; M1 represents asset attribute, including contract transaction price; D1 represents contract content set, including contract terms.

[0047] The step two: through the contract file coding, configuration contract file corresponding project file and financial file includes:

[0048] Read the contract file coding data, through the coding data to find the relevant project file, and the project file coding, form second coding data set S2;

[0049] The second coding data set S2 is expressed as:

[0050]

[0051] Wherein, ID represents the unique identification, including contract name and number; T2 represents time attribute, including project start time, project end time; P2 represents space attribute, including project scope; M2 represents asset attribute, including project budget; D2 represents project file content set, including project plan, project execution, project monitoring;

[0052] Read the contract file coding data, through the coding data to find the relevant financial file, and the financial file coding, form third coding data set S3;

[0053] The third coding data set S3 is expressed as:

[0054]

[0055] Wherein, ID represents the unique identification, including contract name and number; T3 represents time attribute, including the actual use of funds time; P3 represents space attribute, including the use of funds location; M3 represents asset attribute, including financial statements; D3 represents financial file content set, including profitability analysis, operating ability analysis, financial decision;

[0056] Specifically, the contract associated project in the space attribute P1 in the coding information can be used to find out the relevant project file; through the space attribute P1 and the contract associated project and contract transaction price in the asset attribute M1 to find out the relevant financial file;

[0057] The third step of standardizing the encoding of the file to form a standard encoding dataset includes: setting a unique identification ID for the contract file and the associated project file and financial file according to the first encoding dataset S1, the second encoding dataset S2 and the third encoding dataset S3, obtaining the time attribute T, the space attribute P, the asset attribute M and the file content set D of the file through the unique identification ID, and constituting a standard encoding dataset S, wherein .

[0058] Specifically, the contract file and all the project files and financial files associated with the contract file are encoded with a common ID, and in the encoding dataset S, for a contract named “smart factory upgrade project”, all file ID encodings can be identified as “ZNGCxxxx -” as a unique identification, wherein “xxxx” represents the contract number, and these files can be quickly identified as belonging to the specific project.

[0059] After standardizing the encoding of the file, an index file is created, and the standard encoding dataset S is dynamically updated:

[0060] A special encoding index file dataset is set, and in the index file, the ID encoding, file name, file type, main content summary and associated file ID encoding information of each file are listed, and when a new file is generated or the association relationship between files changes, the mapping relationship in the index file is updated in time, and recorded in the standard encoding dataset S.

[0061] Specifically, when a new supplementary agreement is signed, after obtaining the ID encoding, the encoding information of the supplementary agreement is added in the index file, and the association relationship of the contract file and other related files associated therewith is updated, and recorded in the standard encoding dataset under the ID encoding, so that it is convenient to see how the files are connected with each other through encoding.

[0062] The fourth step of reading the standard encoding dataset to obtain file information includes: in the standard encoding dataset S, reading the unique identification ID of the contract file to find the associated file content, automatically displaying other files associated with the unique identification ID, using natural language technology to read the text, and obtaining the file content according to the demand.

[0063] Specifically, when viewing the contract file of the smart factory upgrade project, reading the ID “ZNGCxxxx -”, through the ID in the standard encoding dataset, the links of the project file and the financial file related to the contract are displayed on the interface, which facilitates the user to quickly navigate to the related files, and at the same time, the system also supports the user to search for a group of files associated therewith by inputting the ID or part of the ID;

[0064] Specifically, when text reading is performed by using the natural language technology, the text in the standard coding data set S is preprocessed first, including Chinese word segmentation and part-of-speech tagging, wherein the BIO marking method is used to mark the labels in the natural language, to obtain a natural language question text marking sequence, and a skip-gram model is selected to convert the word vector of the natural language question, each word as a center word is predicted and adjusted for k times to realize more accurate representation, which is expressed as:

[0065]

[0066] wherein P t is a probability, representing the probability of each word appearing as a context word given the center word t, w t is the center word, Sup(w k ) is the context word of the center word, and k is the window size, and P t is maximized through training iteration.

[0067] The objective function is:

[0068]

[0069] wherein L represents the loss, and S represents the standard coding data set.

[0070] The method of machine learning is used for named entity recognition to identify the entity information in the file, such as the signing date, contract term, contract related project, project location, contract transaction price, contract clause, address, contact information, subject of the contract, and breach of contract liability, and to assist in judging the natural language question type, and the convolutional neural network and / or recurrent neural network are used to classify the text to obtain the required text content; for example, the start time and transaction price of the project with ID code “ZNGCxxxx -” are found, and through the natural language technology, the specific content of the start time and transaction price of the project in the contract can be found in the standard coding data set.

[0071] The step five: generating business archive early warning information according to the file information includes: setting early warning conditions for the file about completeness, accuracy, security, compliance and timeliness according to the obtained file content, and giving early warning prompt when abnormal situation occurs.

[0072] Specifically, the early warning information for the file includes:

[0073] The early warning condition for completeness: the file is missing a key page, report or attachment; the early warning information is set: the file “unique ID” is missing a key page, report or attachment, please supplement in time.

[0074] Early warning condition for accuracy: the clause does not conform to the company policy or legal regulations, the data or information is inaccurate; set early warning information: the clause of the file "unique ID" does not conform to the company policy or legal regulations, please check in time, the data or information is inaccurate, please correct in time.

[0075] Early warning condition for security: the file is accessed or tampered by unauthorized access; set early warning information: the file "unique ID" is accessed or tampered by unauthorized access, please check immediately.

[0076] Early warning condition for compliance: the file is not archived or stored as required; early warning information: the file "unique ID" is not archived or stored as required, please handle in time;

[0077] Early warning condition for timeliness: the file or the stage task or report in the file is about to expire or has expired; early warning information: the file "unique ID" is about to expire, please renew or handle in time.

[0078] The application also provides an intelligent business archive management system for realizing the intelligent business archive management method, as shown in Figure 2 The system comprises a reading module, a configuration module, an encoding module, a processing module and an early warning module.

[0079] The reading module reads a contract file in a business archive and encodes the contract file; the configuration module configures a project file and a financial file corresponding to the contract file through the encoding of the contract file; the encoding module standardizes the encoding of the file to form a standard encoding data set; the processing module reads the standard encoding data set to obtain file information; and the early warning module generates business archive early warning information according to the file information.

[0080] In the specification provided herein, a large number of specific details are described. However, it can be understood that the embodiments of the application can be practiced without these specific details. In some examples, well-known methods, structures and techniques are not shown in detail in order not to obscure the understanding of the present specification.

[0081] In addition, those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments means to be within the scope of the application and form different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

Claims

1. An intelligent business file management method, characterized by, The method comprises the following steps: Step one: reading a contract file in a business archive and coding the contract file; comprising: scanning the contract file by optical character recognition technology, identifying text information of the contract file, coding the text information to form a first coded data set S1; Step two: configuring a project file and a financial file corresponding to the contract file through coding of the contract file; comprising: reading coded data of the contract file, searching for a project file related to the contract file through the coded data, and coding the project file to form a second coded data set S2; reading coded data of the contract file, searching for a financial file related to the contract file through the coded data, and coding the financial file to form a third coded data set S3; Step three: standardizing the encoding of the files to form a standard encoding dataset; including: setting a unique identification ID for the contract file and the associated project file and financial file according to the first encoding dataset S1, the second encoding dataset S2 and the third encoding dataset S3, obtaining the time attribute T, the space attribute P, the asset attribute M and the file content set D of the file through the unique identification ID, and constituting a standard encoding dataset S, wherein ; After standardizing coding of the files, an index file is created, and the standard coded data set S is dynamically updated: a special coded index file data set is set, in the index file, ID coding, file name, file type, main content summary and associated file ID coding information of each file are listed, when a new file is generated or the association between files changes, the mapping relationship in the index file is updated in time, and the mapping relationship is recorded in the standard coded data set S; Step four: reading the standard coded data set to obtain file information; comprising: in the standard coded data set S, reading a unique identification ID of a contract file to search for associated file content, automatically displaying other files associated with the unique identification ID, reading text by natural language technology, and obtaining file content according to requirements; Step five: generating business archive early warning information according to the file information.

2. The intelligent business file management method of claim 1, wherein, The first coded data set S1 is represented as: Wherein, ID represents the unique identification, including contract name and number; T1 represents the time attribute, including the signing date and contract period; P1 represents the space attribute, including the contract associated project and project location; M1 represents the asset attribute, including the contract transaction price; D1 represents the contract content set, including the contract terms.

3. The intelligent business archive management method according to claim 1, characterized in that, The second coded data set S2 is represented as: Wherein, ID represents the unique identity, including contract name and number; T2 represents the time attribute, including project start time, project end time; P2 represents the space attribute, including project scope; M2 represents the asset attribute, including project budget; D2 represents the project file content set, including project plan, project execution, project monitoring; The third coded data set S3 is represented as: Wherein, ID represents the unique identification, including contract name and number; T3 represents the time attribute, including the actual use of funds time; P3 represents the space attribute, including the use of funds location; M3 represents the asset attribute, including the financial statements; D3 represents the financial document content set, including profitability analysis, operating ability analysis, financial decision.

4. The intelligent business file management method of claim 1, wherein, The step five: generating business archive early warning information according to the file information comprises: According to the obtained file content, setting early warning conditions of the file on completeness, accuracy, security, compliance and timeliness, and giving a warning prompt when an abnormal condition occurs.

5. An intelligent business file management system, characterized by, The system is used to implement the intelligent business archive management method according to any one of claims 1-4, and comprises a reading module, a configuration module, a coding module, a processing module and a warning module; The reading module reads a contract file in a business archive and codes the contract file; the configuration module configures a project file and a financial file corresponding to the contract file through coding of the contract file; the coding module standardizes coding of the files to form a standard coded data set; the processing module reads the standard coded data set to obtain file information; and the warning module generates business archive early warning information according to the file information.

Citation Information

Patent Citations

  • Coding and displaying method for integrated management of standard files and archives

    CN117349243A

  • Standardized processing method and system for multi-format file

    CN119025480A