Automatic management method and system for tax archives
Through intelligent automatic identification, labeling and indexing technologies, the problems of low automation, slow processing speed, low query efficiency and poor data security of traditional tax archive management systems are solved, and efficient and intelligent tax archive management is achieved.
Patent Information
- Application Number
- CN202510118079.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
Smart Images

Figure CN120106770A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and in particular to an automatic management method for tax files and an automatic management system for tax files. Background Art
[0002] With the continuous development of information technology and the increasing demand for corporate tax management, traditional tax archive management methods have gradually exposed many problems, especially in terms of file storage, query and management efficiency. Most of the existing tax archive management systems rely on manual management, and the organization, classification and archiving of archive content rely on manual experience, resulting in low efficiency and high error rates. In traditional tax archive management systems, the management of corporate tax archives usually adopts paper archives or simple electronic document storage methods. Most of these archives are not effectively labeled and indexed, and relevant information cannot be quickly and accurately located during inquiries. This traditional method not only wastes a lot of manual time and energy, but is also prone to the risk of file loss and damage, and cannot guarantee data security and privacy.
[0003] Although some modern archive management systems have begun to use simple digital technologies, the shortcomings of these systems are still very obvious. For example, many systems still lack efficient automated processing capabilities and rely on manual classification and manual data entry, resulting in slow processing and prone to errors. In addition, the existing systems are relatively simple in the application of labeling and indexing technologies, and cannot achieve comprehensive and intelligent archive classification and retrieval. Especially when the archive types are complex and the content is huge, the system's retrieval efficiency and accuracy cannot meet actual needs.
[0004] In addition, with the increasing diversification of corporate tax archives (such as qualification certificates, tax incentives, asset losses, tax inspections, etc.), a single archive management system is difficult to cope with the diverse management needs of different types of archives, resulting in obvious deficiencies in the flexibility and scalability of the system. Enterprises need to be able to dynamically adjust the system's query rules and index strategies according to different archive types and user needs, but the existing systems generally lack intelligent adaptive adjustment functions and cannot optimize query rules and index structures in real time, further reducing the efficiency of tax archive management.
[0005] Therefore, the traditional tax archive management system has problems such as low automation, slow processing speed, low query efficiency and poor data security. A new technical solution is urgently needed to solve the defects in the existing solution and improve the management efficiency and intelligence level of tax archives by introducing intelligent automatic identification, labeling and indexing technology. Summary of the invention
[0006] The purpose of the embodiments of the present invention is to provide an automatic management method and system for tax archives, so as to at least solve the problems of low automation, slow processing speed, low query efficiency and poor data security in traditional tax archive management systems.
[0007] In order to achieve the above-mentioned purpose, the first aspect of the present invention provides an automatic management method for tax archives, which method includes: collecting tax archive data to be sorted, performing preprocessing on the tax archive data based on a data processing model trained with multivariate data, and obtaining sorted data; performing content recognition on the sorted data, and performing labeling processing on the recognition results; creating an index of the corresponding data based on the label information corresponding to each data; storing the corresponding tax archive data into the database based on the labeled data and the index corresponding to the data, and opening the query rules for the corresponding data.
[0008] Optionally, the collection rules of the tax file data to be sorted include: collecting relevant tax file data in real time from the internal system of the enterprise based on automated tools; wherein the internal system of the enterprise is a financial system, an ERP system and / or a document management system; automatically scanning the collected relevant tax file data to obtain scanned data; and structuring the scanned data according to a preset format as collected data.
[0009] Optionally, the tax file data is preprocessed based on a data processing model trained on multivariate data to obtain sorted data, including: based on the multivariate data training model, the tax file data to be sorted is cleaned, normalized and feature extracted in sequence; the rules for performing content recognition on the sorted data are: based on OCR technology and / or rule-based NLP algorithm, the text, table and image content in the tax file are automatically recognized respectively.
[0010] Optionally, labeling processing is performed on the recognition results, including: performing reasoning on each recognition result based on a pre-trained labeling algorithm to obtain feature information corresponding to each recognition result; matching relevant labels in a predefined label set based on the feature information of each recognition result, and adding at least one label to each corresponding recognition result based on the matching result.
[0011] Optionally, based on the label information corresponding to each data, an index of the corresponding data is created, including: extracting the label information, text content and / or key fields of the tax file to generate index entries; applying an inverted index algorithm to the extracted keywords or tags to record the occurrence position and frequency of each keyword or tag in the document; based on the full-text index, converting the complete text content in the tax file into an indexable byte stream or character stream, and establishing a full-text retrieval index; combining the inverted index and the full-text index to build a multi-level index structure and store it in a database and / or search engine system; partitioning the index structure to distribute the index data in multiple index tables or shards according to document category and / or date range to complete index creation; wherein the method also includes: during the index creation process, hash mapping the pre-calibrated sensitive fields based on the hash algorithm.
[0012] Optionally, the corresponding tax file data is entered into the warehouse based on the labeled data and the index corresponding to the data, including: executing differentiated warehouse entry rules for structured data and unstructured data; wherein, structured data is stored in a relational database, using a standardized table structure, and establishing associations between data through foreign keys; unstructured data is stored in a non-relational database, saved in a binary format, and recording file paths and metadata; during the warehouse entry process, encrypted storage is performed based on the AES encryption algorithm until the warehouse entry is completed; after the warehouse entry is completed, the method also includes: distributing the data to multiple nodes based on a distributed database and sharding technology, and setting up a regular backup mechanism.
[0013] Optionally, the open rules for query rules of corresponding data are: design query interface according to archive type, label and index structure, adopt RESTful API or GraphQL interface specification; parse user request based on query parser, match index field according to query conditions, and generate corresponding query statement; assign corresponding access rights to different users, control data query scope and access level based on user role; execute query operation through SQL query or NoSQL query engine.
[0014] Optionally, the query rules include: collecting user query behavior data based on log analysis tools to record query frequency, query conditions and access patterns; analyzing query logs based on cluster analysis or decision trees to identify common query patterns and high-frequency query fields; automatically generating new query rules and optimization suggestions based on analysis results, adjusting existing index structures or adding new index fields; regularly updating query optimization strategies, and deploying new query rules and index strategies to the query system through automated scripts.
[0015] The second aspect of the present invention provides an automatic management system for tax archives, which includes: a collection unit, which is used to collect tax archive data to be sorted, and pre-process the tax archive data based on a data processing model trained with multivariate data to obtain sorted data; an identification unit, which is used to perform content identification on the sorted data and label the identification results; an index creation unit, which is used to create an index of corresponding data based on the label information corresponding to each data; a management unit, which is used to execute the storage of corresponding tax archive data based on the labeled data and the index corresponding to the data, and open the query rules for the corresponding data.
[0016] On the other hand, the present invention provides a computer-readable storage medium having instructions stored thereon, which, when executed on a computer, enables the computer to execute the above-mentioned automatic tax file management method.
[0017] Through the above technical scheme, the scheme of the present invention significantly improves the efficiency and intelligence level of tax archive management through an intelligent data processing method. First, by collecting the tax archive data to be sorted and pre-processing it using the data processing model trained by multivariate data, the accuracy and consistency of the data are ensured, laying a solid foundation for subsequent processing. Then, through content recognition technology, the key content in the tax archive is automatically identified, and the recognition results are labeled, so that each archive can obtain accurate classification and labeling, which improves the systematization and intelligence of archive management. Subsequently, the index structure of the archive is created based on the label information, so that the data has efficient retrieval performance when stored. Finally, by combining labeled data with indexes, the archival data is automatically stored, and flexible query rules are opened, further optimizing the data access efficiency and query convenience. Overall, the scheme realizes the automatic identification, intelligent classification, accurate storage and efficient retrieval of tax archives, solves the problems of manual dependence, low efficiency and high error rate in traditional management methods, and significantly improves the automation, intelligence and security of archive management.
[0018] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:
[0020] Figure 1 It is a flowchart of the steps of the automatic management method of tax files provided by one embodiment of the present invention;
[0021] Figure 2It is a system structure diagram of an automatic management system for tax archives provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0022] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.
[0023] Figure 1 FIG. 1 is a flow chart of a method for automatically managing tax files provided by an embodiment of the present invention. Figure 1 As shown, an embodiment of the present invention provides an automatic management method for tax files, the method comprising:
[0024] Step S10: Collect tax file data to be sorted, and perform preprocessing on the tax file data based on a data processing model trained with multivariate data to obtain sorted data.
[0025] Specifically, the collection rules of the tax file data to be sorted include: collecting relevant tax file data from the internal system of the enterprise in real time based on automated tools; wherein the internal system of the enterprise is a financial system, an ERP system and / or a document management system; automatically scanning the collected relevant tax file data to obtain scanned data; and structuring the scanned data according to a preset format as collected data.
[0026] Furthermore, the tax file data is preprocessed based on the data processing model trained with multivariate data to obtain sorted data, including: based on the multivariate data training model, the tax file data to be sorted is cleaned, normalized and feature extracted in sequence.
[0027] In an embodiment of the present invention, the tax file data collection link uses an automated tool to collect relevant tax file data from the internal system of the enterprise in real time. These internal systems include financial systems, ERP systems, and document management systems, etc., which store a large number of tax-related documents and data. In order to ensure the timeliness and comprehensiveness of the data, the collection rules ensure that the system can obtain the latest tax file data from these internal systems in real time or regularly, and process the collected data according to a preset format. In particular, automatic scanning technology is applied to the document scanning and recognition process. For paper documents or scanned images, the system will automatically convert them into digital data, and further structure the scanned data to generate data records in a standard format. This process ensures that the collected tax file data can be easily input into the subsequent processing and management system.
[0028] After data collection is completed, the next step is to pre-process the data using a processing model based on multivariate data training to ensure the quality and applicability of the data. The data processing model is based on multivariate data training and combines machine learning, deep learning and other technologies. Through the learning of a large amount of historical data, the model can effectively identify and process tax file data of various formats and types. The main steps of data pre-processing include:
[0029] 1) Data cleaning: This step automatically identifies and removes invalid, duplicate or erroneous data through algorithms. The system can detect and remove any incomplete, abnormal or irrelevant data for tax file management, such as incomplete files, format errors or garbled characters, etc., to ensure the accuracy of subsequent processing steps.
[0030] 2) Data normalization: This step standardizes the data format so that the archive data from different sources can be unified into a standardized format. This is because the tax archive data stored in different systems (such as financial systems, ERP systems, etc.) may have different codes, date formats, field definitions, etc. Normalization can make all data consistent, thus facilitating subsequent classification and processing.
[0031] 3) Feature extraction: Feature extraction technology analyzes and extracts key fields in archival data. These fields may include tax numbers, tax inspection records, license numbers, tax incentives, etc. Through data mining and natural language processing (NLP) technology, the system can automatically extract key information from the text, structure the data and provide support for subsequent classification and labeling.
[0032] Through this series of preprocessing steps, the resulting "organized data" not only meets the requirements of structured storage, but also provides a basis for subsequent intelligent management and retrieval. Throughout the process, data cleaning, normalization, and feature extraction are combined with machine learning and rule-driven algorithms to ensure the efficiency and accuracy of data processing.
[0033] Based on the scheme of the present invention, tax file data is collected from the internal system of the enterprise through automated tools, which solves the inefficiency and error problems of traditional manual collection and greatly improves the timeliness and comprehensiveness of the data. Secondly, the processing model based on multivariate data training can automatically perform data cleaning, normalization and feature extraction, avoiding omissions and deviations in manual operations. The pre-processed data is highly structured and standardized, laying a solid foundation for subsequent automatic recognition, labeling and index creation.
[0034] In addition, the training of the data processing model enables the system to adapt to tax files of different formats and types, solving the limitation of the existing system that cannot handle complex data formats and types. Through this technical solution, the management of tax files becomes more efficient and accurate, and can quickly adapt to various changes and needs in the management of corporate tax files, ensure the security and integrity of data, and provide strong support for subsequent intelligent query and analysis.
[0035] Step S20: performing content recognition on the sorted data, and performing labeling processing on the recognition results.
[0036] Specifically, the rules for performing content recognition on the collated data are: based on OCR technology and / or rule-based NLP algorithm, the text, table and image contents in the tax files are automatically recognized respectively.
[0037] Furthermore, labeling processing is performed on the recognition results, including: performing reasoning on each recognition result based on a pre-trained labeling algorithm to obtain feature information corresponding to each recognition result; matching relevant labels in a predefined label set based on the feature information of each recognition result, and adding at least one label to each corresponding recognition result based on the matching result.
[0038] In the embodiment of the present invention, the sorted tax file data involves various forms of content, such as text, tables, pictures, etc. In order to accurately and efficiently extract valuable information, these data are first automatically identified. Specifically, different types of content in the tax file are identified by combining OCR technology (optical character recognition technology) and rule-based natural language processing (NLP) algorithms.
[0039] OCR technology is used to recognize images or scanned documents in tax files, especially for processing scanned images of paper documents. Through image preprocessing (such as denoising, binarization, tilt correction, etc.), the OCR engine can accurately extract text content. For documents in different languages and fonts, the system will select the appropriate OCR model as needed, and use deep learning for training to further improve recognition accuracy.
[0040] NLP algorithms are mainly used to identify text information in tax files, including ordinary text, table content, and various tax terms. Through technologies such as word segmentation, syntactic analysis, and semantic understanding, the system can identify key information in the file, such as tax ID, license number, tax inspection records, etc. In addition, NLP can also handle complex tax terms and legal clauses to ensure a comprehensive understanding of the file content.
[0041] Furthermore, after the content recognition is completed, in order to facilitate the classification, query and management of the archive data, the system will label the recognition results. The specific process of labeling includes:
[0042] Reasoning and feature extraction: Based on the pre-trained labeling algorithm, the system performs reasoning analysis on each piece of data identified. Through the pre-trained model, the system can extract relevant feature information from the recognition results, such as the subject, keywords, date, numbers, etc. of the text. For image content, the text information extracted by OCR can also be further analyzed to obtain relevant feature data.
[0043] Based on the feature information extracted from the recognition results, the system will match in the predefined label set. These label sets include various labels for tax files, such as "qualification certificate", "tax incentives", "asset loss", "tax inspection", etc. The content of the label set can be customized according to the needs of specific enterprises. Through the algorithm, the system will automatically match the relevant labels for each file based on the identified content features to ensure that each file can be accurately labeled.
[0044] Each tax file can be matched with multiple tags to ensure that the system can fully capture all aspects of the file. For example, a qualification certificate file containing tax inspection records may be labeled with both "qualification certificate" and "tax inspection". This not only helps managers quickly find relevant files, but also provides more refined classification and retrieval functions.
[0045] Based on the scheme of the present invention, the combination of OCR technology and NLP algorithm ensures accurate recognition of various file formats and contents, not only limited to traditional text files, but also documents in complex formats such as scanned copies and pictures. This technology greatly reduces the workload of manual input and inspection, avoids human errors, and improves the speed and accuracy of data processing. Labeling automatically tags each file with multiple tags through a predefined tag set, making the classification of files more refined. Unlike traditional single classification, labeling can support more complex file management needs, such as querying and managing according to multiple dimensions at the same time, improving the efficiency and flexibility of file retrieval. The use of reasoning and feature extraction technology ensures that even complex text content and image data can be effectively analyzed and classified by machine learning algorithms. This enables the system to adapt to various types and formats of tax files, which not only improves the automation level of file management, but also improves the scalability and flexibility of the system. Finally, the query rule design based on tags greatly simplifies the file query process, and users can perform fast and accurate retrieval through multiple tag combinations. This improvement reduces the time for enterprises to find files in daily tax management and improves the efficiency and accuracy of tax decision-making.
[0046] Step S30: creating an index of the corresponding data based on the tag information corresponding to each data.
[0047] Specifically, the label information, text content and / or key fields of the tax file are extracted to generate index entries; an inverted index algorithm is applied to the extracted keywords or tags to record the occurrence position and frequency of each keyword or tag in the document; based on the full-text index, the complete text content in the tax file is converted into an indexable byte stream or character stream, and a full-text retrieval index is established; combining the inverted index and the full-text index, a multi-level index structure is constructed and stored in a database and / or search engine system; the index structure is partitioned, and the index data is distributed in multiple index tables or shards according to the document category and / or date range to complete the index creation; wherein the method also includes: during the index creation process, hash mapping is performed on pre-calibrated sensitive fields based on a hash algorithm.
[0048] In an embodiment of the present invention, the system generates index entries by extracting tag information, text content and key fields in tax files. Tax files contain various types of data, such as tax numbers, qualification certificate numbers, tax inspection records, etc. The system automatically identifies and extracts these key fields, and classifies them based on the tag information to form structured data items. Through natural language processing (NLP) technology and rule-driven extraction methods, the system can accurately identify important information in various types of documents and automatically generate index entries based on the tags of each file. Each index entry contains keywords, tag information and its specific location in the document.
[0049] Furthermore, the extracted keywords or tags are processed using an inverted index algorithm. An inverted index is a data structure commonly used in information retrieval systems that can efficiently record the location and frequency of each keyword or tag in a document. Specifically, the system will traverse each tax file and extract important words, tags, and their locations. The inverted index records the specific location where each keyword or tag appears, so that relevant documents can be quickly located when querying. This method greatly improves retrieval efficiency, especially when faced with a large number of tax files, relevant files can be quickly found by keywords or tags.
[0050] In the process of creating the index, the system will also convert the complete text content in the tax file into an indexable byte stream or character stream based on full-text indexing technology. Full-text indexing can analyze and index the entire document content and map each word, phrase or paragraph of the file to the database. By converting the document into a byte stream, the system can not only process structured text data, but also effectively process various types of unstructured text, such as PDF files, text content in scanned images, etc. These contents will be converted into a standardized format and stored in the index to facilitate subsequent retrieval and analysis.
[0051] Combining the inverted index and the full-text index, the system will build a multi-level index structure. This index structure can not only improve the efficiency of the query, but also support more complex retrieval requirements. For example, the system can organize the index according to multiple dimensions such as the document's tag information, keywords, date, and file type. The multi-level index structure allows the system to select the corresponding index level according to different query conditions when processing query requests, thereby speeding up the query response time. In addition, the system uses a database and / or search engine system (such as Elasticsearch) to store these index structures to ensure high data availability and distributed query capabilities.
[0052] In order to improve the scalability and query performance of the index, the system adopts index partitioning and sharding strategies during the index creation process. According to the document category (such as qualification certificates, tax incentives, asset losses, etc.) and attributes such as date range, the system will distribute the index data in multiple index tables or shards. For example, the index is partitioned by year or archive category so that each query operation can quickly locate the corresponding partition, thereby reducing the search scope during the query. The sharding strategy can also improve the parallel processing capability of the query, especially when the amount of data is large, which can significantly improve the performance of the system.
[0053] During the index creation process, in order to enhance the security of sensitive data, the system will also use hash algorithms to hash pre-calibrated sensitive fields. After sensitive fields such as tax numbers and license numbers are hashed, even indexed data no longer exposes the original data content. The hash algorithm protects privacy and avoids the risk of data leakage by converting sensitive data into irreversible hash values. Hash mapping ensures that sensitive information will not be exposed to unauthorized users during data storage and retrieval, while also preserving the integrity of indexing and retrieval functions.
[0054] Based on the scheme of the present invention, this scheme significantly improves the storage, retrieval and management efficiency of tax archive data. The combination of inverted index and full-text index makes the query operation more efficient, especially in large-scale archive management systems, which can quickly locate relevant documents and greatly shorten the query time. The multi-level index structure based on tag information and key fields not only enhances the scalability of the index, but also supports complex multi-dimensional queries, allowing users to accurately find the required archives according to different query conditions. In addition, index partitioning and sharding technology improves the scalability and performance of the system under massive data, allowing the system to handle larger-scale tax archives. Finally, the hash mapping method for processing sensitive data effectively protects data privacy and security and prevents potential security risks. The combination of these technologies not only solves the query efficiency and data security problems in traditional tax archive management systems, but also provides enterprises with an efficient, secure and flexible tax archive management solution.
[0055] Step S40: Based on the labeled data and the index corresponding to the data, the corresponding tax file data is stored in the database, and the query rules for the corresponding data are opened.
[0056] Specifically, the corresponding tax file data is entered into the warehouse based on the labeled data and the index corresponding to the data, including: executing differentiated warehouse entry rules for structured data and unstructured data; wherein, structured data is stored in a relational database, using a standardized table structure, and establishing associations between data through foreign keys; unstructured data is stored in a non-relational database, saved in a binary format, and recording file paths and metadata; during the warehouse entry process, encrypted storage is performed based on the AES encryption algorithm until the warehouse entry is completed; after the warehouse entry is completed, the method also includes: distributing the data to multiple nodes based on a distributed database and sharding technology, and setting up a regular backup mechanism.
[0057] In the embodiment of the present invention, after the labeling of the tax file is completed, the system performs a storage operation on the tax file data based on the labeling data and the corresponding index information, ensuring that the data can be stored efficiently and securely, and providing support for subsequent query and analysis. This process includes adopting differentiated storage rules for structured data and unstructured data, so as to maximize the advantages of various data storage systems.
[0058] For structured data, this data usually includes core information in tax files, such as tax numbers, license numbers, tax inspection records, asset losses, etc. After being labeled, this information can generate standardized fields and relationships. The system stores this structured data in a relational database, such as MySQL, PostgreSQL, or DAMO database. The data is stored in a standardized table structure, and each table contains clear field definitions to ensure data consistency and integrity. At the same time, tables are associated with each other through foreign key relationships, so that different types of tax file data can be effectively integrated and referenced. The use of foreign key associations can ensure the referential integrity of the data and avoid the generation of isolated data, thereby achieving more efficient data query and management.
[0059] For unstructured data, which usually includes scanned images, PDF files, or other document formats, it is difficult to store them in traditional structured ways. This type of data is converted into a standard format through OCR (optical character recognition) technology or other forms of preprocessing, and then stored in a non-relational database such as MongoDB, Cassandra, or a distributed file storage system (such as HDFS). Unstructured data is usually saved in binary format (such as PDF, image files), and the corresponding file path and metadata (such as file size, type, upload time, etc.) are recorded in the database. This storage method can handle large-scale document storage needs and facilitate subsequent access and management.
[0060] During the entire storage process, all data is encrypted and stored using the AES encryption algorithm to ensure data security and privacy. AES encryption is a symmetric encryption algorithm that encrypts data before storage to prevent data from being illegally accessed or tampered with during storage. The encrypted data will be stored in the database or file system, ensuring the integrity and security of the data and complying with data protection and privacy regulations. Once the data is stored, the system will also use distributed databases and sharding technology to store the data in a distributed manner. The data will be stored on multiple nodes in a decentralized manner, and the sharding technology will be used to divide a large amount of data into smaller units for more efficient management and access. Distributed storage not only improves the data processing capability, but also enhances the fault tolerance and scalability of the system, and can adapt to tax archive data of different sizes. In addition, the system will also set up a regular backup mechanism to ensure that the data can be restored in the event of hardware failure or other emergencies. The backup mechanism usually includes full backup and incremental backup to ensure data persistence and reliability. Regular backup operations can effectively avoid data loss and ensure high availability of the system.
[0061] Based on the solution of the present invention, the efficiency, security and scalability of tax archive data storage are improved by adopting differentiated warehousing rules and advanced encryption technology. The standardized storage of structured data makes the relationship between data clearer and the query efficiency higher. Unstructured data is stored using the flexibility of non-relational databases, which solves the problem that traditional relational databases are difficult to handle large-scale document data. The use of the AES encryption algorithm ensures the privacy and security of the data and prevents the risk of data leakage during the storage process. Through distributed databases and sharding technology, the system can handle large-scale data storage needs and improve the reliability and query efficiency of data storage. In addition, the regular backup mechanism ensures that data can be restored under any circumstances, greatly improving the system's disaster recovery capabilities and data security.
[0062] Furthermore, the open rules for query rules of corresponding data are: design query interface according to archive type, label and index structure, and adopt RESTful API or GraphQL interface specification; parse user request based on query parser, match index field according to query conditions, and generate corresponding query statement; assign corresponding access rights to different users, control data query scope and access level based on user role; execute query operation through SQL query or NoSQL query engine.
[0063] Furthermore, the query rules include: collecting user query behavior data based on log analysis tools to record query frequency, query conditions and access patterns; analyzing query logs based on cluster analysis or decision trees to identify common query patterns and high-frequency query fields; automatically generating new query rules and optimization suggestions based on analysis results, adjusting existing index structures or adding new index fields; regularly updating query optimization strategies, and deploying new query rules and index strategies to the query system through automated scripts.
[0064] In the embodiment of the present invention, the scheme of the present invention designs a flexible and efficient query interface based on the file type, label and index structure, so that users can quickly obtain relevant tax information according to different needs. The open design of the query rule adopts the RESTful API or GraphQL interface specification, which enables the query operation to not only meet the needs of the system internally, but also facilitate the integration of external systems.
[0065] The design of the query interface is first based on the type, label and index structure of the archive, and organizes this data in the interface specification. As the mainstream Web interface specifications today, the RESTful API interface and the GraphQL interface provide a concise and efficient way to handle user requests. The RESTful API uses standard HTTP methods (GET, POST, PUT, DELETE, etc.) to interact with resources, and specifies query conditions through URLs and query parameters; GraphQL provides a more flexible query method, and users can accurately specify the returned data fields according to their needs, and can request multiple different data resources at one time, avoiding the problem of multiple requests. Both the RESTful API and the GraphQL interface can efficiently support data interaction between different users and systems.
[0066] The user's query request will be parsed by the query parser, which determines the specific fields and conditions of the query based on the query conditions entered by the user. Then, the parser will map the query conditions to the index fields in the database and generate the corresponding query statements. This process ensures the flexibility and efficiency of the query, especially when processing multi-condition combination queries, the system can quickly locate relevant data through indexes and generate the optimal query statement through efficient query algorithms.
[0067] In order to ensure the security and privacy of the system, the query operation also includes user rights management. The system assigns corresponding access rights according to the user's role, and controls the scope and access level of their data query based on the user's role. For example, ordinary users can only query tax files within their authority, while administrators can access all file data. Through the role-based access control (RBAC) mechanism, the system can ensure that users of different roles can only query and operate data within their authorized scope, avoiding the risk of sensitive data leakage.
[0068] After generating the query statement and completing the permission verification, the system will perform the actual query operation through SQL query or NoSQL query engine. For structured data, the system executes SQL query through relational database (such as MySQL, PostgreSQL); for unstructured data, it uses NoSQL database (such as MongoDB, Elasticsearch) to execute the query operation. Based on the different database types, the system can select the most appropriate query method to improve the efficiency and accuracy of the query.
[0069] In order to continuously improve the query performance and intelligence level of the system, this technical solution also adds query log analysis and adaptive query optimization mechanisms. Through log analysis tools, the system can collect and record user query behavior data, including query frequency, query conditions, access patterns and other information. These log data will provide support for subsequent query optimization and rule generation.
[0070] By performing cluster analysis or decision tree analysis on query logs, the system can identify common query patterns and high-frequency query fields. For example, certain file categories may be frequently queried, or certain fields (such as tax numbers, license numbers, etc.) often appear in query conditions. Based on these analysis results, the system can automatically generate new query rules and optimization suggestions. For example, new indexes can be created for high-frequency query fields, or the existing index structure can be adjusted according to the hot search conditions of the query to further improve query efficiency.
[0071] The system automatically adjusts query rules and index structures based on the analysis results, and deploys new query optimization strategies and index strategies to the query system through automated scripts to ensure that query performance can be continuously optimized. The automated optimization process reduces manual intervention and improves the system's adaptability and flexibility in changing usage environments. Regularly updating query optimization strategies helps the system remain efficient and responsive when facing different query requirements and changes in data volume.
[0072] Based on the scheme of the present invention, through these technical means, the tax file query system can realize efficient and secure query services. The query design based on RESTful API and GraphQL interface enables the system to support internal queries and be easily integrated with other systems. The efficient combination of query parser and index structure ensures the accuracy and rapid response of queries. At the same time, the user authority management mechanism ensures the privacy and security of data and prevents unauthorized data access. The introduction of query optimization and adaptive adjustment enables the system to continuously optimize query rules and index structure as user query behavior changes, thereby maintaining the improvement of query performance. Log analysis and cluster analysis technology further improve the system's sensitivity to query patterns and self-learning ability, thereby continuously improving query efficiency and accuracy. This series of technical measures effectively solves the performance bottleneck in traditional query systems, ensuring that the system can efficiently and securely support the management and query needs of corporate tax files.
[0073] Figure 2 1 is a system structure diagram of an automatic management system for tax files provided by one embodiment of the present invention. Figure 2As shown, an embodiment of the present invention provides an automatic management system for tax archives, and the system includes: a collection unit, which is used to collect tax archive data to be sorted, and pre-process the tax archive data based on a data processing model trained with multivariate data to obtain sorted data; an identification unit, which is used to perform content identification on the sorted data and label the identification results; an index creation unit, which is used to create an index of corresponding data based on the label information corresponding to each data; a management unit, which is used to execute the storage of corresponding tax archive data based on the labeled data and the index corresponding to the data, and open the query rules of the corresponding data.
[0074] The embodiment of the present invention further provides a computer-readable storage medium, on which instructions are stored, which, when executed on a computer, enable the computer to execute the above-mentioned automatic management method of tax files.
[0075] Those skilled in the art will understand that all or part of the steps in the method for implementing the above-mentioned embodiments can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including several instructions for making a single-chip microcomputer, a chip or a processor (processor) perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0076] The optional embodiments of the present invention are described in detail above in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the technical concept of the embodiments of the present invention, the technical scheme of the embodiments of the present invention can be subjected to a variety of simple modifications, and these simple modifications all belong to the protection scope of the embodiments of the present invention. It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the embodiments of the present invention will not further describe various possible combinations.
[0077] In addition, various embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the embodiments of the present invention, they should also be regarded as the contents disclosed in the embodiments of the present invention.
Claims
1. A method for automatically managing tax files, characterized in that: The method comprises: Collecting tax file data to be sorted, and performing preprocessing on the tax file data based on a data processing model trained with multivariate data to obtain sorted data; Performing content recognition on the collated data, and performing labeling processing on the recognition results; Based on the label information corresponding to each data, create an index for the corresponding data; Based on the labeled data and the corresponding index of the data, the corresponding tax archive data is stored in the database and the query rules for the corresponding data are opened.
2. The method according to claim 1, characterized in that The collection rules of the tax file data to be sorted include: Based on automated tools, relevant tax file data is collected from the internal system of the enterprise in real time; The enterprise internal system is a financial system, an ERP system and / or a document management system; Automatically scan the collected relevant tax file data to obtain scanned data; The scanned data is structured according to a preset format as collected data.
3. The method according to claim 1, characterized in that The tax file data is preprocessed based on the data processing model trained with multivariate data to obtain sorted data, including: Based on the multivariate data training model, the tax file data to be sorted is cleaned, normalized and feature extracted in sequence; The rules for performing content identification on the collated data are: Based on OCR technology and / or rule-based NLP algorithms, the text, table and image contents in tax files are automatically recognized.
4. The method according to claim 1, wherein the labeling process is performed on the recognition result, comprising: Perform inference on each recognition result based on the pre-trained labeling algorithm to obtain feature information corresponding to each recognition result; Based on the feature information of each recognition result, relevant tags are matched in a predefined tag set, and at least one tag is added to each corresponding recognition result based on the matching result.
5. The method according to claim 1, creating an index of corresponding data based on the tag information corresponding to each data, comprising: Extract label information, text content and / or key fields of tax files to generate index entries; Apply an inverted index algorithm to the extracted keywords or tags to record the location and frequency of each keyword or tag in the document; Based on full-text indexing, the complete text content in the tax archives is converted into an indexable byte stream or character stream, and a full-text search index is established; Combine inverted index and full-text index to build a multi-level index structure and store it in the database and / or search engine system; The index structure is partitioned and the index data is distributed in multiple index tables or shards according to document categories and / or date ranges to complete index creation. The method further comprises: During the index creation process, the pre-calibrated sensitive fields are hash mapped based on the hash algorithm.
6. The method according to claim 1, characterized in that Based on the labeled data and the corresponding index of the data, the corresponding tax file data is stored in the database, including: Differentiated warehousing rules are implemented for structured data and unstructured data; Structured data is stored in a relational database, using a standardized table structure and establishing associations between data through foreign keys; Unstructured data is stored in a non-relational database in binary format, and the file path and metadata are recorded; During the storage process, encrypted storage is performed based on the AES encryption algorithm until the storage is completed; After the storage is completed, the method further includes: Based on distributed database and sharding technology, data is distributed to multiple nodes, and a regular backup mechanism is set up.
7. According to the method of claim 1, the open rule of the query rule corresponding to the data is: Design query interfaces based on archive types, tags, and index structures, using RESTful API or GraphQL interface specifications; Parse user requests based on the query parser, match index fields according to query conditions, and generate corresponding query statements; Assign corresponding access rights to different users and control data query scope and access level based on user roles; Execute query operations through SQL query or NoSQL query engine.
8. The method according to claim 1, wherein the query rule comprises: Collect user query behavior data based on log analysis tools to record query frequency, query conditions, and access patterns; Analyze query logs based on cluster analysis or decision trees to identify common query patterns and high-frequency query fields; Based on the analysis results, new query rules and optimization suggestions are automatically generated to adjust the existing index structure or add new index fields; Regularly update query optimization strategies and deploy new query rules and index strategies to the query system through automated scripts.
9. An automatic management system for tax files, characterized in that: The system comprises: A collection unit, used for collecting tax file data to be sorted, performing preprocessing on the tax file data based on a data processing model trained with multivariate data, and obtaining sorted data; An identification unit is configured to perform content identification on the collated data and label the identification result; An index creation unit, used to create an index of corresponding data based on label information corresponding to each data; The management unit is used to execute the storage of corresponding tax file data based on the labeled data and the index corresponding to the data, and to open the query rules of the corresponding data.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the automatic management method for tax files as described in any one of claims 1 to 8.
Citation Information
Cited By
Data asset sharing method and system
CN120930913A