A digital electricity invoice management method and cloud platform based on AI model
By deploying AI models and virtual printers on electronic devices and combining with cloud platforms, seamless integration and unified processing of existing business systems are achieved, high cost and complexity problems of traditional electronic invoice is solved, and a low-cost and easy-to-implement digital invoice management solution is provided.
Patent Information
- Application Number
- CN202510234954.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The traditional electronic invoice issuance method requires enterprises to invest a lot of money in business system upgrades, hardware updates, personnel training, etc., which is expensive and lacks flexibility, making it difficult to meet rapidly changing tax requirements and compatibility of different business systems.
Using the digital invoice management method based on the AI model, by deploying AI models, virtual printers and synchronization tools on electronic devices, combined with the cloud platform, seamless integration and unified processing of existing business systems are achieved, types of bills are automatically identified, and file parsing and invoice is used to generate digital invoices and print them.
It reduces the cost of system upgrade and transformation, improves compatibility and automation, adapts to changes in tax policies, reduces hardware and labor costs, solves network security restrictions, and provides a low-cost and easy-to-implement digital invoice management solution.
Smart Images

Figure CN119722362B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of cloud data processing, and in particular to a digital electricity invoice management method and cloud platform based on an AI model. Background Art
[0002] With the rapid development of my country's economy and the deepening of its digital transformation, electronic invoices, as a more efficient and environmentally friendly form of billing, are gradually replacing traditional paper invoices. Digital invoices not only improve invoicing efficiency and reduce paper usage, but also effectively prevent counterfeit invoices and improve tax management. However, during this transformation, many companies face the challenge of quickly and cost-effectively upgrading their existing systems to adapt to new invoicing requirements.
[0003] Traditional electronic invoicing methods typically rely on existing business systems. However, these systems are typically designed to meet daily operational needs and lack the ability to directly issue electronic invoices. To implement electronic invoicing, companies typically need to take the following steps: 1. System upgrade: Companies must entrust the original system developer or a third-party software company to upgrade their existing systems. This includes adding an electronic invoice module, modifying data structures, and adjusting business processes. This process typically requires extensive development work, encompassing system analysis, design, coding, and testing. 2. Hardware upgrade: To support electronic invoice issuance and management, companies must purchase new servers, storage devices, and other hardware to meet the increased data processing and storage requirements. 3. Tax control equipment integration: Companies typically need to purchase dedicated tax control equipment (such as tax control disks) and integrate them with their business systems. This not only increases hardware costs but also requires specialized technical personnel to maintain and manage the equipment. 4. Network environment modification: For security reasons, business systems in some specialized industries (such as finance and healthcare) typically do not allow direct access to the external network. To implement electronic invoicing, these companies need to transform their network environments and add security isolation measures, further increasing complexity and costs. 5. Compliance Adjustment: As tax policies change, electronic invoicing systems need to be constantly updated and adjusted to meet the latest compliance requirements. Companies need to continuously invest resources to maintain and update the system.
[0004] In summary, traditional electronic invoicing methods require businesses to invest substantial capital in a short period of time for business system upgrades, hardware upgrades, and personnel training. This prohibitive cost for small and medium-sized enterprises (SMEs) makes widespread adoption difficult. Furthermore, because each company's business system is unique, customized upgrades struggle to achieve economies of scale and reduce costs for individual businesses. Therefore, while existing business system upgrades can achieve basic functionality, their high costs, complex implementation, and lack of flexibility make them inadequate for today's rapidly changing business environment and increasingly stringent tax requirements. Summary of the Invention
[0005] The embodiments of the present application provide a digital invoice management method and cloud platform based on an AI model, which are used to effectively reduce the cost of transforming traditional business systems and effectively solve the problems of complex implementation process and lack of flexibility of traditional electronic invoice issuance methods.
[0006] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:
[0007] In a first aspect, a digital electricity invoice management method based on an AI model is provided, which is applied to an electronic device, wherein the electronic device is deployed with an AI model, a business platform, a virtual printer, and a synchronization tool, and the electronic device is connected to a physical printer. The method includes:
[0008] Responding to a billing signal from the business platform, obtaining an electronic business bill generated by the business platform;
[0009] Identify the electronic business bill using the AI model to obtain a domain type of the electronic business bill;
[0010] Converting the electronic business bill into a file format using the virtual printer to obtain a converted electronic business bill, and storing the converted electronic business bill in a corresponding folder, wherein the folder is generated based on the domain type of the electronic business bill;
[0011] In response to an event trigger signal of the folder, uploading the converted electronic business invoice to a cloud platform via the synchronization tool, wherein the synchronization tool is used to monitor the folder, and the cloud platform is used to parse the converted electronic business invoice to obtain invoice information, execute an invoicing process based on the invoice information, generate a digital invoice and a callback instruction, and send the digital invoice and callback instruction to the synchronization tool, so that the synchronization tool generates a printing instruction based on the callback instruction, and the printing instruction is used to control the physical printer;
[0012] The printing instruction and the digital invoice are sent to the physical printer, so that the physical printer prints the digital invoice according to the printing instruction.
[0013] In a possible implementation of the first aspect, identifying the electronic business bill using the AI model to obtain the domain type of the electronic business bill includes:
[0014] Extracting features of the electronic business bill to obtain text features and metadata features of the electronic business bill;
[0015] The text features and the metadata features are input into the AI model to obtain the domain type of the electronic business bill, wherein the domain type includes the HIS domain and the ERP domain.
[0016] In another possible implementation of the first aspect, the step of generating a folder includes:
[0017] Obtaining a first timestamp of generation of the electronic business bill;
[0018] When the domain type is the HIS domain, obtaining a business code of a business platform for generating the electronic business ticket, and using a first preset naming rule to generate a first name for the folder according to the first timestamp;
[0019] Generate a first path of the folder based on a preset first basic path, the first timestamp and the business code;
[0020] Based on the first name and the first path, create a folder, and create a HIS domain main folder and at least one HIS domain subfolder in the folder;
[0021] Allocating a first permission to the HIS domain main folder and a second permission to at least one of the HIS domain subfolders, and recording a second path of the HIS domain main folder and a second timestamp of creation of the HIS domain main folder;
[0022] Creating a first JSON file in the HIS domain main folder and generating a log, wherein the first JSON file includes the second timestamp and the domain type, and the log includes the second path, the first permission, and the second permission;
[0023] When the domain type is the ERP domain, obtaining a business class code of the business platform that generates the electronic business bill, and using a second preset naming rule to generate a second name for the folder according to the first timestamp;
[0024] generating a third path of the folder based on a preset second basic path, the first timestamp, and the business class code;
[0025] Based on the second name and the third path, a folder is created, and an ERP domain main folder and at least one ERP domain subfolder are created in the folder, wherein the number of the ERP domain subfolders is the same as the number of the HIS domain subfolders;
[0026] Allocating a third permission to the ERP domain main folder and a fourth permission to at least one of the ERP domain subfolders, and recording a fourth path of the ERP domain main folder and a third timestamp of creation of the HIS domain main folder;
[0027] A second json file is created in the ERP domain main folder, and a log is generated, wherein the second json file includes the third timestamp and the domain type, and the log includes the fourth path, the third permission, and the fourth permission.
[0028] In another possible implementation of the first aspect, uploading the converted electronic business document to a cloud platform through the synchronization tool includes:
[0029] The synchronization tool uses the converted electronic business bill as a file to be synchronized, and identifies the file type and file size of the file to be synchronized;
[0030] The synchronization tool determines a corresponding compression algorithm according to the file type, and compresses the file to be synchronized using the compression algorithm to obtain a compressed file;
[0031] The synchronization tool identifies the sensitive fields of the compressed file and encrypts the sensitive fields to obtain an encrypted data packet;
[0032] The synchronization tool establishes an HTTPS connection with the cloud platform and uploads the encrypted data packet segments to the cloud platform.
[0033] In another possible implementation of the first aspect, the synchronization tool compresses the to-be-synchronized file using the compression algorithm to obtain a compressed file, including:
[0034] The synchronization tool creates a memory buffer and initializes the compression algorithm;
[0035] The synchronization tool reads the to-be-synchronized file into memory in blocks, and compresses each data block of the to-be-synchronized file using the compression algorithm;
[0036] The synchronization tool writes all the compressed data blocks into a new file to obtain a compressed file.
[0037] In another possible implementation of the first aspect, the synchronization tool encrypts the sensitive field to obtain an encrypted data packet, including:
[0038] For any sensitive field, the synchronization tool obtains the position of the sensitive field in the file to be synchronized and generates a random initialization vector;
[0039] The synchronization tool uses a preset encryption algorithm to replace the sensitive field with the random initialization vector to obtain an encrypted sensitive field;
[0040] The synchronization tool generates encryption metadata according to the random initialization vector and the position of the sensitive field in the file to be synchronized;
[0041] The synchronization tool inserts the encrypted sensitive field into the corresponding position of the file to be synchronized to obtain an encrypted file;
[0042] The synchronization tool packages the encrypted metadata and the encrypted file into a container file to obtain an encrypted data packet.
[0043] In another possible implementation of the first aspect, the synchronization tool establishes an HTTPS connection with the cloud platform and uploads the encrypted data packet fragments to the cloud platform, including:
[0044] The synchronization tool calculates a first file checksum of the encrypted data packet and generates upload metadata, wherein the upload metadata includes the first name, the file size, and the first file checksum, or the upload metadata includes the second name, the file size, and the first file checksum;
[0045] The synchronization tool establishes an HTTPS connection with the cloud platform;
[0046] The synchronization tool divides the encrypted data packet into a plurality of fixed-size data blocks, uploads the data blocks to the cloud platform one by one, and records the fixed-size data blocks uploaded to the cloud platform;
[0047] After the synchronization tool uploads all the fixed-size data blocks to the cloud platform, the synchronization tool sends the upload metadata to the cloud platform.
[0048] In a second aspect, the present application provides a digital invoice management method based on an AI model, which is applied to the above-mentioned cloud platform, wherein the cloud platform is connected to an electronic device, and the electronic device is connected to a physical printer, and the method is characterized in that it includes:
[0049] In response to receiving one of the fixed-size data blocks transmitted by the electronic device, storing the fixed-size data block in a preset buffer until all the fixed-size data blocks are received;
[0050] reorganizing all the fixed-size data blocks in the buffer to obtain a reorganized file;
[0051] Calculating a second file checksum of the reorganized file;
[0052] Receive and parse the uploaded metadata to obtain the first file verification code and the location of the sensitive field;
[0053] comparing the second file checksum with the first file checksum;
[0054] After the comparison is successful, the encrypted data packet is unpacked to obtain the encrypted metadata and the encrypted file, and the field to be decrypted is located according to the position of the sensitive field;
[0055] Decrypting the field to be decrypted using a preset decryption algorithm and a preset security key to obtain the sensitive field corresponding to the field to be decrypted, and obtaining a decrypted compressed file;
[0056] Decompressing the compressed file using a preset decompression algorithm to obtain an electronic business receipt;
[0057] Parsing the electronic business bill to obtain bill information, executing the billing process according to the bill information, and generating a digital invoice and a callback instruction;
[0058] The digital invoice and callback instruction are sent to the electronic device so that the electronic device generates a printing instruction according to the callback instruction, and the printing instruction and the digital invoice are sent to the physical printer, which is used to print the digital invoice according to the printing instruction.
[0059] In a possible implementation of the second aspect, the invoicing process includes:
[0060] Call the preset tax interface, obtain the invoice code and invoice number based on the bill information, and generate an electronic signature;
[0061] Fill the invoice code, invoice number and electronic signature into the preset standard template to generate a digital invoice.
[0062] In a third aspect, the present application provides a cloud platform, which is applied to the digital electricity invoice management method based on the AI model of the first aspect, and the digital electricity invoice management method based on the AI model of the second aspect, including:
[0063] a communication module for communicating with an electronic device; and
[0064] The data processing module is connected to the communication module and is used to execute the digital electricity invoice management method based on the AI model of the second aspect.
[0065] The above technical solution, firstly, by deploying the AI model, business platform, virtual printer, and synchronization tools on electronic devices, seamlessly integrates existing business systems, eliminating the need for large-scale upgrades and modifications to existing enterprise systems, significantly reducing implementation costs and complexity. This effectively avoids the tedious process of entrusting legacy system developers or third-party software companies with system upgrades, saving significant development effort and associated costs. Secondly, the use of a virtual printer effectively resolves the issue of inconsistent output formats across different business systems, enabling unified processing of electronic business invoices without the need for developing separate interfaces for each system, significantly improving compatibility and scalability. The introduction of an AI model further enhances intelligence, enabling automatic identification of the field type of electronic business invoices, enabling precise classification and processing, and increasing the level of automation across the entire process. Furthermore, this technical solution utilizes a cloud-based platform for file parsing and invoicing, reducing the burden on local systems and enabling better adaptation to tax policy changes. Updates can be performed in the cloud, eliminating the need for frequent adjustments to internal enterprise systems, significantly reducing the cost and difficulty of compliance adjustments. In response to the network security restrictions faced by some special industries, this technical solution realizes the interaction of internal and external network data through synchronization tools and folder monitoring mechanisms, which not only ensures data security but also meets the invoicing needs, and solves the problem of complex network environment transformation required in traditional methods. In terms of hardware investment, this technical solution mainly relies on the company's existing electronic equipment and physical printers. There is no need to purchase additional dedicated tax control equipment or a large amount of new hardware, which significantly reduces hardware costs. At the same time, due to the high degree of automation of the entire process, the need for personnel training is greatly reduced, labor costs are reduced, and work efficiency is improved. In general, this technical solution effectively overcomes the problems of high cost, complex implementation, and lack of flexibility in traditional electronic invoicing methods, and provides enterprises, especially small and medium-sized enterprises, with a low-cost, easy-to-implement and well-adaptable digital invoice management solution.
[0066] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 A schematic diagram of a process for issuing electronic invoices for an enterprise provided in an embodiment of the present application;
[0068] Figure 2 A schematic diagram of the first process of a digital electricity invoice management method based on an AI model provided in an embodiment of the present application;
[0069] Figure 3 A second flow chart of a method for managing digital electricity invoices based on an AI model provided in an embodiment of the present application;
[0070] Figure 4 A schematic diagram of a folder creation process provided in an embodiment of the present application;
[0071] Figure 5 A third flow chart of a method for managing digital electricity invoices based on an AI model provided in an embodiment of the present application;
[0072] Figure 6 A schematic diagram of the structure of a cloud platform provided in an embodiment of the present application. DETAILED DESCRIPTION
[0073] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific implementation methods described herein are only used to illustrate and explain the embodiments of the present application and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0074] It should be noted that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back, etc.), such directional indications are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0075] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0076] To facilitate understanding of the solution, please refer to Figure 1 , Figure 1 A schematic diagram of a process for issuing an electronic invoice for an enterprise provided in an embodiment of the present application is shown. The process is as follows:
[0077] 1. The enterprise submits an application to the finance and taxation management agency through its business system. The application content includes the required invoice type, quantity and other information; 2. The finance and taxation management agency reviews the enterprise's application and approves a certain number of invoicing shares based on the enterprise's tax credit, business scale and other factors, and sends the approved shares electronically to the enterprise's business system; 3. When the enterprise actually needs to issue an invoice, it uses the invoice issuance module in the business system and enters relevant transaction information, such as product name, quantity, amount, etc. The business system automatically generates an electronic invoice and consumes the previously applied invoicing share; 4. Print the issued invoice. If a paper version is required, the business system will send the electronic invoice to the printer; 5. Make an invoice declaration and report it to the finance and taxation management agency. The enterprise regularly (usually monthly) summarizes the invoice information issued and declares the invoice through the business system or special declaration software. The declaration content includes detailed information such as the number of invoices issued, the amount, tax amount, etc., and the declaration information is uploaded to the system of the finance and taxation management agency.
[0078] From this process, we can see that the invoice issuance module, by being integrated into the enterprise's business system, achieves a seamless connection between business data and invoice issuance.
[0079] An enterprise's business system is usually deployed on the enterprise's central control host. Therefore, based on this process, the embodiment of the present application proposes a digital electricity invoice management method based on an AI model, which can effectively solve the problem of high cost of integrating traditional invoice issuance modules with business systems.
[0080] Specifically, Figure 2 The following schematically shows a flow chart of a method for managing digital electricity invoices based on an AI model according to an embodiment of the present application. Figure 2 and Figure 3 As shown, an embodiment of the present application provides a digital invoice management method based on an AI model, which is applied to an electronic device. The electronic device is deployed with an AI model, a business platform, a virtual printer and a synchronization tool. The electronic device is connected to a physical printer. The method may include the following steps.
[0081] S110. Responding to an invoicing signal from the business platform, obtaining an electronic business invoice generated by the business platform;
[0082] S120. Identify the electronic business bill using the AI model to obtain the domain type of the electronic business bill.
[0083] S130: Convert the electronic business bill into a file format using a virtual printer to obtain a converted electronic business bill, and store the converted electronic business bill in a corresponding folder, wherein the folder is generated based on the domain type of the electronic business bill;
[0084] S140. In response to an event trigger signal from the folder, the converted electronic business invoice is uploaded to the cloud platform via a synchronization tool. The synchronization tool is used to monitor the folder, and the cloud platform is used to parse the converted electronic business invoice to obtain invoice information, execute an invoicing process based on the invoice information, generate a digital invoice and a callback instruction, and send the digital invoice and the callback instruction to the synchronization tool, so that the synchronization tool generates a print instruction based on the callback instruction, and the print instruction is used to control a physical printer.
[0085] S150: Send the printing instruction and the digital invoice to the physical printer, so that the physical printer prints the digital invoice according to the printing instruction.
[0086] In this embodiment, the electronic device can be a tablet computer, desktop computer, laptop computer, handheld computer, wearable device, notebook computer, ultra-mobile personal computer (UMPC), netbook computer, or other device with a processor. Of course, the electronic device can also be a server. The specific form of the electronic device is not particularly limited in this embodiment of the application.
[0087] The electronic device first responds to the invoicing signal sent by the business platform. The invoicing signal is triggered by the user clicking the "Issue Invoice" button on the business platform interface, or it can be automatically triggered by the business system according to preset rules.
[0088] Upon receiving the invoicing signal, the electronic device immediately establishes a data connection with the business platform and retrieves the relevant electronic business receipt through a predefined interface or API call. This electronic business receipt may include the transaction amount, product or service description, buyer and seller information, and more. During the retrieval process, the electronic device performs preliminary data verification to ensure its integrity and accuracy. If any data is missing or formatted incorrectly, feedback is immediately provided to the business platform, requesting the correct data be re-provided.
[0089] In this embodiment, in order to improve processing efficiency, the electronic device can adopt a batch processing method to obtain multiple electronic business invoices to be issued at one time. The obtained invoice data is temporarily stored in the memory or local storage of the electronic device.
[0090] After receiving electronic business invoices, an AI model is used to identify and classify them. The AI model is a trained deep learning model, such as a convolutional neural network (CNN) or a recurrent neural network (RNN). It analyzes the structure, content, and characteristics of electronic business invoices to determine their domain type. In this solution, the primary domain types identified are HIS (Hospital Information System) and ERP (Enterprise Resource Planning System). During the identification process, the AI model considers multiple factors, such as the format of the invoice, the fields included, and the terminology used. For example, if a invoice contains a large amount of medical-related terminology or patient information, the model identifies it as an HIS type; if a invoice primarily contains information related to business operations, such as inventory, order, or financial data, it is identified as an ERP type. The use of AI models significantly improves recognition accuracy and efficiency, enabling rapid processing of large numbers of electronic business invoices from various sources. This allows the electronic device to adopt appropriate processing strategies based on the different types of invoices, improving the intelligence and efficiency of the entire process.
[0091] A virtual printer is a software-simulated printing device that intercepts the data stream sent to the printer and converts it into files of other formats. In this embodiment, the virtual printer first receives the original electronic business invoices from the business platform. These invoices may exist in a variety of different formats, such as PDF, Word documents, Excel spreadsheets, etc. The virtual printer can convert these invoices in different formats into a standardized format according to preset rules, usually PDF or a specific image format. During the conversion process, the virtual printer retains all the key information of the original invoice, while ensuring that the converted file has good readability and compatibility. After the conversion is completed, the electronic device will automatically create or select the corresponding folder in the storage space of the electronic device based on the field type (HIS or ERP) previously identified by the AI model. Specifically, Figure 4 As shown, a folder path can be generated based on the field type of the electronic business invoice, and folders can be created based on the folder path. The created folders can be organized by date, field type, or other rules. The converted electronic business invoices are then stored in the corresponding folders, effectively improving the processing capabilities of invoices from different sources and formats, providing a unified data format for subsequent cloud upload and processing, and facilitating data traceability.
[0092] like Figure 2As shown, the electronic business invoices in the folder will be transferred to the synchronization tool for processing. Specifically, the synchronization tool is used to continuously monitor the folder created in step S130. When an event is detected in the folder (such as a new file being stored, an existing file being modified or deleted), the synchronization tool will immediately trigger the upload operation. During the upload process, the synchronization tool will establish a secure connection with the cloud platform and can use encrypted transmission protocols such as HTTPS to protect data security. After the file upload is completed, the cloud platform receives the converted electronic business invoice and will immediately start the file parsing process. The parsing process can use OCR technology to extract text information from PDFs or images, or use a preset parsing algorithm to process structured data. The parsed invoice information includes transaction details, amount, date, buyer and seller information, etc.
[0093] The cloud platform then executes the invoicing process based on the parsed invoice information, combined with pre-set rules and templates. This process involves generating a digital invoice that meets tax requirements and creating a callback instruction containing the processing results and subsequent instructions. The generated digital invoice and callback instruction are then sent back to the synchronization tool. Upon receiving the digital invoice and callback instruction, the synchronization tool generates a print instruction based on the callback instruction. The print instruction contains detailed parameters for printing the digital invoice, such as paper size, print quality, and number of copies. This ultimately achieves a seamless integration of local data and cloud processing, significantly improving the efficiency and accuracy of invoice processing while leveraging cloud resources to reduce the burden on local systems.
[0094] Finally, the electronic device sends the print instructions and digital invoice data generated in step S140 to the connected physical printer. This transmission process typically occurs via a standard printing protocol, such as the Internet Printing Protocol (IPP), or a vendor-specific print driver. Once the physical printer receives this data, it sets printing parameters, such as page layout, print quality, and color mode, based on the received print instructions. The printer then begins the printing process, converting the digital invoice content into a physical paper document. Upon completion, the printer sends a print completion status message to the electronic device, allowing the business platform to update the invoice status. This embodiment ensures that digital invoices can be converted to physical form, meeting the needs of scenarios requiring paper invoices. Furthermore, by precisely controlling the printing process, the quality and compliance of the printed invoices are guaranteed. The entire process is highly automated, reducing human intervention, improving efficiency, and lowering error rates.
[0095] In summary, compared with traditional business system transformation, this embodiment has the following effects: by deploying AI models, virtual printers and synchronization tools, seamless integration with existing business platforms is achieved without the need for major modifications to the original system, effectively solving the problem that traditional methods require large-scale transformation of existing systems, and reducing system upgrade and transformation costs; using virtual printers instead of dedicated tax control equipment reduces hardware costs. At the same time, the complex bill parsing and invoicing processes are processed by the cloud platform, reducing the requirements for local hardware performance; uploading data to the cloud platform through synchronization tools solves the problem that some special industries (such as finance and medical care) do not allow the system to directly access the external network due to security considerations; the AI model recognizes bills in different fields (HIS and ERP) and stores them in corresponding folders, realizing the effective integration of different business systems; the cloud platform is responsible for executing the invoicing process and generating digital invoices, which can more easily adapt to changes in tax policies and perform unified updates without the need for enterprises to maintain them themselves; the embodiment of the present application can be quickly deployed and is suitable for enterprises of different sizes and types, overcoming the problem that customized transformation in traditional methods is difficult to form economies of scale; since the solution mainly relies on mature AI technology and cloud services, the implementation risk is greatly reduced compared to the comprehensive transformation of existing systems; the automatic recognition of AI models, automatic conversion of virtual printers, automatic processing of cloud platforms and other features greatly improve the efficiency of the entire invoicing process; the modular design (AI model, virtual printer, synchronization tool, cloud platform) makes the business system have good scalability and adaptability, and can better cope with changes in the business environment.
[0096] In general, this embodiment effectively solves the problems of high cost, complex implementation, lack of flexibility and other issues in traditional methods through the combination of intelligence, automation and cloud services, and provides a more economical, efficient, easy to implement and well-adaptable solution, which is especially suitable for enterprises that want to quickly and cost-effectively upgrade to a digital invoice system.
[0097] First, this embodiment achieves seamless integration of existing business systems by deploying AI models, business platforms, virtual printers, and synchronization tools on electronic devices, eliminating the need for large-scale upgrades and modifications to the company's existing systems, significantly reducing implementation costs and complexity. This effectively avoids the tedious process of entrusting the original system developer or third-party software company to perform system upgrades, as is the case with traditional methods, saving a significant amount of development work and related expenses. Second, the use of virtual printers effectively resolves the issue of inconsistent output formats across different business systems, enabling unified processing of electronic business invoices without the need to develop separate interfaces for each system, significantly improving compatibility and scalability. The introduction of AI models further enhances intelligence, enabling automatic identification of the domain type of electronic business invoices, thereby achieving accurate classification and processing, and improving the automation level of the entire process. Furthermore, this technical solution utilizes a cloud platform for file parsing and invoicing process processing, which not only reduces the burden on local systems but also enables better adaptation to changes in tax policies. Updates only require cloud-based updates, eliminating the need for frequent adjustments to the company's internal systems, significantly reducing the cost and difficulty of compliance adjustments. In response to the network security restrictions faced by some special industries, this technical solution realizes the interaction of internal and external network data through synchronization tools and folder monitoring mechanisms, which not only ensures data security but also meets the invoicing needs, and solves the problem of complex network environment transformation required in traditional methods. In terms of hardware investment, this technical solution mainly relies on the company's existing electronic equipment and physical printers. There is no need to purchase additional dedicated tax control equipment or a large amount of new hardware, which significantly reduces hardware costs. At the same time, due to the high degree of automation of the entire process, the need for personnel training is greatly reduced, labor costs are reduced, and work efficiency is improved. In general, this technical solution effectively overcomes the problems of high cost, complex implementation, and lack of flexibility in traditional electronic invoicing methods, and provides enterprises, especially small and medium-sized enterprises, with a low-cost, easy-to-implement and well-adaptable digital invoice management solution.
[0098] In one implementation of this embodiment, identifying an electronic business bill using an AI model to obtain the domain type of the electronic business bill includes the following steps:
[0099] S210: Extract features of the electronic business bill to obtain text features and metadata features of the electronic business bill;
[0100] S220. Input text features and metadata features into the AI model to obtain the domain type of the electronic business bill, where the domain type includes the HIS domain and the ERP domain.
[0101] In this embodiment, the feature extraction process is divided into two main aspects: extraction of text features and metadata features.
[0102] First, text feature extraction primarily involves keyword identification and text structure analysis. For keyword identification, natural language processing (NLP) techniques, such as word frequency-independent frequency (TF-IDF) or word embedding algorithms, can be used to extract key words from electronic business documents. For example, in the HIS (hospital information system) domain, key words might be medical-related terms like "patient," "visit," and "prescription." For ERP (enterprise resource planning) domains, key words might be business terms like "order," "supplier," and "invoice." Furthermore, the position, frequency, and contextual relationships of words within a document can be identified. For example, if the word "patient" appears frequently in a document and frequently co-occurs with words like "diagnosis" and "treatment," then the document belongs to the HIS domain.
[0103] Text structure analysis involves identifying document titles, paragraph divisions, and table structures. Specific formatting patterns can be identified using basic formatting rules. For example, documents in the HIS domain contain specific structures such as patient information tables and diagnostic reports, while documents in the ERP domain contain different structural features such as order details and financial statements.
[0104] Secondly, metadata feature extraction is used to extract non-content information carried by the document itself. This includes the document's generation time, generation, and related IDs. This information is usually stored in the document's attributes or header. For example, the document's creation date and last modification time can be extracted. These timestamps are used to indicate the document's timeliness and importance. Generated information, such as identifiers such as "HIS-v3.2" or "ERP-2022," directly indicates the document's source and provides clues for domain classification. For example, the format and structure of related IDs, such as patient IDs or order numbers, also indicate the document's domain attributes.
[0105] The extracted features are converted into numerical vectors or other formats suitable for processing by machine learning models. For example, keywords are converted into word frequency vectors, document structure is encoded as binary features, and timestamps are converted into the number of days relative to a base date.
[0106] Afterwards, the extracted features are input into the AI model to derive the domain type of the electronic business bill.
[0107] Specifically, the text features and metadata features extracted in step S210 are first preprocessed and normalized. This includes operations such as feature scaling (e.g., normalizing all numerical features to a range of 0-1), handling missing values (e.g., filling with the mean or median), and feature encoding (e.g., converting categorical features to one-hot encoding). This ensures that all features are at the same scale, which facilitates AI model learning and prediction.
[0108] Next, the preprocessed features are fed into a pre-trained AI model. This AI model can be a deep learning model, such as a multi-layer perceptron (MLP) or a long short-term memory (LSTM). The AI model consists of multiple hidden layers. Each layer applies a nonlinear transformation to the input features, extracting progressively higher-level abstract features. For example, the first layer may learn simple word combination patterns, while subsequent layers may capture more complex semantic relationships. The final layer is a softmax layer, which outputs a probability distribution for each category (HIS and ERP).
[0109] The model is trained on a large number of labeled electronic business invoice samples. During the training phase, the model continuously adjusts its internal parameters using a backpropagation algorithm to minimize the discrepancy between its predictions and the true labels. Once trained, the model is capable of classifying new, unseen electronic business invoices.
[0110] In practice, when a new electronic business invoice needs to be classified, its text and metadata features are fed into a trained AI model. The AI model then calculates the probability of the invoice belonging to the HIS domain or the ERP domain. For example, the AI model outputs a probability of 0.85 for the HIS domain and a probability of 0.15 for the ERP domain. The category with the highest probability is selected as the final classification result. In this example, the electronic business invoice is classified as belonging to the HIS domain.
[0111] This implementation uses in-depth feature extraction and advanced AI models to accurately identify the domain type of electronic business invoices, whether they are HIS or ERP. This automated classification not only greatly improves processing efficiency and reduces manual intervention, but also provides a reliable foundation for subsequent business processes.
[0112] In one implementation of this embodiment, the step of generating a folder includes:
[0113] S301, obtaining a first timestamp of the electronic business bill;
[0114] In the case where the domain type is the HIS domain, obtaining a business code of a business platform for generating an electronic business bill, and using a first preset naming rule to generate a first name for the folder according to the first timestamp;
[0115] S302: Generate a first path of a folder based on a preset first basic path, a first timestamp, and a business code;
[0116] S303: Create a folder based on the first name and the first path, and create a HIS domain main folder and at least one HIS domain subfolder in the folder;
[0117] S304: Allocate a first permission for the HIS domain main folder and a second permission for at least one HIS domain subfolder, and record a second path of the HIS domain main folder and a second timestamp of creation of the HIS domain main folder;
[0118] S305. Create a first json file in the HIS domain main folder and generate a log, where the first json file includes the second timestamp and the domain type, and the log includes the second path, the first permission, and the second permission;
[0119] S306: When the domain type is the ERP domain, obtain the business class code of the business platform for generating the electronic business invoice, and use the second preset naming rule to generate a second name for the folder based on the first timestamp;
[0120] S307: Generate a third path of the folder based on the preset second basic path, the first timestamp, and the business class code;
[0121] S308. Create a folder based on the second name and the third path, and create an ERP domain main folder and at least one ERP domain subfolder in the folder, wherein the number of ERP domain subfolders is the same as the number of HIS domain subfolders;
[0122] S309: Allocate a third permission to the ERP domain main folder and a fourth permission to at least one ERP domain subfolder, and record a fourth path of the ERP domain main folder and a third timestamp of creation of the HIS domain main folder;
[0123] S310. Create a second json file in the ERP domain main folder and generate a log. The second json file includes a third timestamp and domain type. The log includes a fourth path, a third permission, and a fourth permission.
[0124] First, obtain the first timestamp of the electronic business document generation. In this embodiment, the timestamp is a number accurate to the millisecond, representing the exact time the document was generated. For example, "20240419153045123" represents 15:30:45:123 on April 19, 2024. The timestamp can be obtained by calling the system time function. The timestamp is not only used for folder naming but also serves as part of the unique identifier of the document, ensuring that each document has a unique timestamp.
[0125] Next, if the domain type is confirmed to be a HIS domain, obtain the business code of the business platform that generated the electronic business invoice. The business code is a predefined string used to identify different medical business types, such as "OPD" for outpatient departments and "IPD" for inpatient departments. The business code can be obtained by querying the database or reading the configuration file.
[0126] Then, a first preset naming rule is used to generate a first name for the folder based on the first timestamp. This naming rule can use the following format: "HIS_[business code]_[timestamp]." This folder name clearly indicates that the folder belongs to the HIS domain; the business code can be used to identify folders of different business types; and the use of the timestamp ensures that each folder name is unique, preventing conflicts even if multiple folders are created in the same second.
[0127] For example, if the timestamp is 20240419153045123 and the business code is OPD, the generated folder name is: "HIS_OPD_20240419153045123".
[0128] A first path of the folder is generated based on a preset first basic path, a first timestamp, and a business code to determine the exact location of the newly created folder in the file system.
[0129] First, the preset first basic path is a directory pre-defined in the system, which is stored in the configuration file or hard-coded in the program. For example, the basic path can be: " / data / HIS_documents / ". Next, a time-related directory structure is created using the first timestamp. The timestamp can be converted into the form of year, month, and day. For example, if the date corresponding to the timestamp is May 1, 2023, the following path structure can be generated: " / data / HIS_documents / 2023 / 05 / 01 / ". Then, the business code is added to the path. If the business code is "OPD", the complete path becomes: " / data / HIS_documents / 2023 / 05 / 01 / OPD / ". Finally, the folder name generated in step S301 is added to the end of the path to form the final first path.
[0130] Based on the first name and the first path generated in steps S301 and S302 , a folder is created, and a HIS domain main folder and at least one HIS domain subfolder are created in the folder to physically create a directory structure in the file system.
[0131] First, check if the first path already exists. If not, recursively create the entire directory structure. This can be done using the directory creation functions provided by your programming language, such as the Files.createDirectories() method in Java. These functions automatically create all non-existent parent directories in the path.
[0132] After creating the main directory, create a HIS domain main folder in it. This main folder can be named "HIS_Main" or other names to store the main documents and metadata related to the HIS domain. Next, create at least one HIS domain subfolder. Subfolders correspond to different document types. For example, the subfolder Patient_Records is used to store patient records; the subfolder converted is used to store converted electronic business invoices; the subfolder Billing is used to store billing information, etc. It can be set according to actual conditions, and this application does not limit this.
[0133] Assign primary permissions to the HIS domain master folder and secondary permissions to at least one HIS domain subfolder. Record the secondary path to the HIS domain master folder and the secondary timestamp of the HIS domain master folder's creation. First, assign primary permissions to the HIS domain master folder. This permission setting utilizes an operational access control list (ACL) or similar mechanism. For example, the following permissions can be set: Owner (typically an administrator or HIS user): Read, Write, Execute; Group (perhaps a medical staff group): Read, Execute; Others: No permissions. This permission setting ensures that only authorized personnel can access and modify the contents of the HIS master folder.
[0134] Next, assign secondary permissions to each HIS domain subfolder. Secondary permissions can be set based on the specific purpose of the subfolder. For example, set strict permissions on the Patient_Records folder, allowing access only to specific medical personnel; and limit access to the Billing folder to members of the Finance department.
[0135] Then, the second path of the HIS domain main folder will be recorded. The second path is different from the first path generated in step S302, which specifically refers to the location of the main folder. For example, if the first path is:
[0136] " / data / HIS_documents / 2023 / 05 / 01 / OPD / HIS_OPD_20240419153045123 / ";
[0137] Then the second path can be:
[0138] " / data / HIS_documents / 2023 / 05 / 01 / OPD / HIS_OPD_20240419153045123 / HIS_Main / ".
[0139] Recording the second path can facilitate subsequent quick location and access to the HIS main folder without traversing the entire directory structure.
[0140] Finally, the second timestamp of the creation of the HIS domain master folder is recorded, which accurately records the actual creation time of the master folder.
[0141] Create the first JSON file in the main folder of the HIS domain and generate a log. Specifically, create a JSON format file in the main folder of the HIS domain, which can be named "metadata.json". The JSON file contains two key pieces of information: the second timestamp and the domain type. For example, the file content may be {"timestamp": "20240419153046789","domain": "HIS"}. At the same time, generate a log file to record more detailed information, including the second path, the first permission (main folder permission) and the second permission (subfolder permission). The log file can be named "XX.txt".
[0142] When the domain type is ERP (Enterprise Resource Planning), you need to obtain the business class code of the business platform that generates the electronic business invoice. The business class code is a more detailed business identifier, for example, "SALES_ORDER" represents a sales order, or "PURCHASE_ORDER" represents a purchase order. Then, using a second preset naming rule, a secondary folder name is generated based on the first timestamp. This naming rule differs from that used in the HIS domain and can adopt the format "ERP_Business Class Code_Timestamp." So, if the first timestamp is "20240419153045123" and the business class code is "SALES_ORDER," the secondary folder name could be "ERP_SALES_ORDER_20240419153045123."
[0143] The third path for the folder is generated based on the preset second base path, the first timestamp, and the business class code. The second base path is the preset storage root directory for the ERP domain, similar to the first base path for the HIS domain, and can be " / data / ERP / ." Combining the second base path with the previously generated second name creates the complete third path. For example, if the second base path is " / data / ERP / " and the second name is "ERP_SALES_ORDER_20240419153045123," the generated third path is " / data / ERP / ERP_SALES_ORDER_20240419153045123 / ."
[0144] Based on the generated second naming and third path, create a folder structure for the ERP domain. First, create a main folder, and then create an ERP domain main folder and at least one ERP domain subfolder in it. The ERP domain main folder can be named "ERP_MAIN", and the number of subfolders needs to be consistent with the subfolders in the HIS domain to ensure the correspondence between the data structures of the two systems. For example, you can create subfolders such as "CUSTOMER_INFO", "ORDER_DETAILS", and "FINANCIAL_RECORDS". The number should be consistent with the subfolders in the HIS domain. This will facilitate comparison and integration with the data in the HIS system when receiving electronic invoice data from both domains at the same time.
[0145] Assign permissions to the ERP master folder and subfolders. Similar to HIS permissions, you can set permissions as follows: Owner (usually an administrator or ERP user): Read, Write, Execute; Group (which can be a company personnel group): Read, Execute; Others: No permissions. This permission setting ensures that only authorized personnel can access and modify the contents of the ERP master folder.
[0146] Next, assign secondary permissions to each ERP domain subfolder. Secondary permissions can be set based on the specific purpose of the subfolder. For example, the PURCHASE_ORDER folder can be set with strict permissions, allowing access only to specific purchasing personnel; the SALES_ORDER folder can be accessed only by members of the sales department. This permission setting ensures the security of ERP data and prevents unauthorized access or modification.
[0147] At the same time, record the fourth path of the ERP domain main folder. This path is the absolute path of the ERP main folder, for example, " / data / ERP / ERP_SALES_ORDER_20240419153045123 / ERP_MAIN / ". You also need to record the third timestamp of the ERP main folder creation. The third timestamp reflects the actual creation time of the ERP folder.
[0148] A second JSON file is created in the ERP domain's main folder. The JSON file contains two key pieces of information: the third timestamp and the domain type. For example, the file might contain {"timestamp": "20240419153047456", "domain":"ERP"}. A new log file is also generated to record detailed ERP-related information, including the fourth path, third permissions (ERP main folder permissions), and fourth permissions (ERP subfolder permissions). The log file can be named "XX.txt."
[0149] The folder generation and management solution of this embodiment provides a unified, orderly, and secure framework for data storage in two different areas: HIS and ERP. By using timestamps and business codes / category codes to generate unique folder names, the uniqueness and traceability of data storage are ensured. The hierarchical folder structure (main folders and subfolders) allows different types of data to be clearly classified and stored, facilitating subsequent data processing and analysis. The allocation of permissions ensures data security and prevents unauthorized access. The creation of JSON files and logs provides the system with important metadata, which assists in system maintenance, troubleshooting, and data analysis. In particular, by maintaining consistency in the number of subfolders in the HIS and ERP systems, the foundation is laid for data integration and comparative analysis between the two systems.
[0150] In one implementation of this embodiment, uploading the converted electronic business invoice to the cloud platform through a synchronization tool includes the following steps:
[0151] S410: The synchronization tool uses the converted electronic business invoice as a file to be synchronized, and identifies the file type and file size of the file to be synchronized;
[0152] S420: The synchronization tool determines a corresponding compression algorithm based on the file type, and compresses the file to be synchronized using the compression algorithm to obtain a compressed file.
[0153] S430, the synchronization tool identifies sensitive fields of the compressed file and encrypts the sensitive fields to obtain an encrypted data packet;
[0154] S440. The synchronization tool establishes an HTTPS connection with the cloud platform and uploads the encrypted data packet segments to the cloud platform.
[0155] The synchronization tool first uses the converted electronic business invoice as the file to be synchronized, and identifies the file type and file size of the synchronized file. Identifying the file type is achieved by analyzing the file extension and content. The file extension can provide preliminary type information, such as .pdf for PDF files and .xlsx for Excel files. In order to more accurately determine the file type, this embodiment also needs to check the magic number of the file. The magic number is a string of specific bytes at the beginning of the file that is used to identify the file type. For example, the magic number of a PDF file is "%PDF-". This can be implemented using Python's magic library.
[0156] The file size can be obtained through the file system API. The file size value is returned in bytes. For example, a 1MB PDF file returns 1048576 bytes.
[0157] For example, suppose the electronic business invoice to be synchronized is a 5MB PDF file. The synchronization tool first reads the first few bytes of the file and confirms that the magic number is "%PDF-," thus identifying it as a PDF file. It then uses the file system API to obtain the file size, which is 5,242,880 bytes. This information is recorded for subsequent processing steps.
[0158] Based on the file type identified in step S410, the synchronization tool will select an appropriate compression algorithm. The selection process is usually based on preset rules or configuration files. Different types of files are suitable for different compression algorithms because their data structures and redundancy characteristics are different. For example, for text files (such as .txt, .xml, .json), you can choose the GZIP general compression algorithm; for image files, if lossless compression is required, you can choose the PNG compression algorithm; if lossy compression is allowed, you can choose JPEG compression. For file formats that have already been compressed (such as .pdf, .docx), you can choose a lightweight compression algorithm or skip the compression step directly, because these files are usually already optimized, and re-compression may not be effective or may even increase the file size.
[0159] The degree of compression is affected by the compression level, which typically ranges from 1 (fastest compression) to 9 (best compression). Choosing the right compression level is a trade-off between file size reduction and compression time. For a 5MB PDF file, choosing a medium compression level (such as 5) might reduce the file size to about 3-4MB, and the compression time might range from a few hundred milliseconds to several seconds.
[0160] The compression process is implemented using specialized compression libraries, such as zlib. After compression, the synchronization tool generates a new compressed file, typically with an extension indicating the compression algorithm, such as .pdf.gz. By compressing the synchronized files, the file size is significantly reduced, thereby reducing the time and bandwidth required for network transmission. This improves data transmission efficiency for files containing large amounts of repetitive content or structured data.
[0161] After compression is complete, the synchronization tool identifies sensitive fields in the compressed file and encrypts them. Sensitive fields can include personal identification information (such as name, ID number, phone number), financial information (such as bank account number, credit card information), medical information, etc. The definition of sensitive fields is usually stored in configuration files or hard-coded in the program. Regular expressions can be used to identify sensitive fields. In the HIS field, sensitive fields can include patient personal information, diagnosis results, etc.; in the ERP field, sensitive fields can include financial data, contract amounts, etc.
[0162] Assuming that during the parsing process, the synchronization tool identifies an ID card number and a bank card number, these sensitive information needs to be encrypted. The encryption algorithm can be symmetric encryption (such as AES) or asymmetric encryption (such as RSA). This embodiment uses a symmetric encryption algorithm. Taking AES-256 encryption as an example, the process is as follows:
[0163] 1. Generate a random 256-bit key. 2. Use this key to encrypt sensitive information. For example, "310****34" is encrypted to "7Xt2p9Q3Rf7yLm1nO8Bv4W==". 3. Replace the original sensitive information with the encrypted text. 4. Transmit the encryption key and share it with the recipient (cloud platform) via a secure key exchange protocol (such as Diffie-Hellman).
[0164] An encrypted data packet typically consists of two parts: the encrypted file and encrypted metadata. This embodiment significantly improves data security. Even if the data is intercepted during transmission, an attacker without the correct decryption key cannot read sensitive information. Furthermore, by encrypting only sensitive fields rather than the entire file, security is guaranteed while also balancing processing efficiency and avoiding unnecessary computational overhead.
[0165] The synchronization tool establishes a secure connection with the cloud platform and uploads the encrypted data packets in segments. First, the synchronization tool establishes an HTTPS connection with the cloud platform. HTTPS is the abbreviation for HTTP over SSL / TLS, which provides encryption and authentication for data transmission through the SSL / TLS protocol. The process of establishing an HTTPS connection is as follows: 1. The client (synchronization tool) sends a "Hello Client" message to the server (cloud platform), which contains information such as the TLS version and encryption algorithm supported by the client. 2. The server responds with a "Hello Server" message, selects the TLS version and encryption algorithm to use, and sends its digital certificate. 3. The client verifies the server's certificate to ensure that it is indeed communicating with the intended cloud platform. 4. The client generates a pre-master key, encrypts it with the server's public key, and sends it to the server. 5. Both parties use the pre-master key to generate a session key, and all subsequent communications are encrypted using this session key.
[0166] After establishing a secure connection, the synchronization tool begins uploading the encrypted data package. Considering that the data package may be large (even after compression), a multi-part upload strategy is usually adopted. The multi-part upload process is as follows: 1. Divide the encrypted data package into blocks of a fixed size, for example, 5MB per block. 2. Calculate the checksum (such as MD5 or SHA256) for each shard for subsequent integrity verification. 3. Initialize the upload task and obtain the upload ID. 4. Upload each shard in turn. Each upload contains the shard sequence number, shard data, and checksum. 5. After all shards are uploaded, a request to complete the upload is sent, which contains information about all shards. 6. The cloud platform verifies the integrity of all shards. If the verification passes, the shards are assembled into a complete file.
[0167] After the upload is complete, the synchronization tool will send a request to complete the upload, informing the server (cloud platform) that all shards have been uploaded and the file can be assembled.
[0168] This implementation prioritizes the use of an appropriate compression algorithm, significantly reducing file size and improving transmission efficiency. The identification and encryption of sensitive information ensures data security during transmission and storage, effectively protecting user privacy. Finally, by establishing a secure HTTPS connection and implementing a multi-segment upload strategy, data transmission security is ensured while also improving the reliability and efficiency of large file transfers.
[0169] In one implementation of this embodiment, the synchronization tool uses a compression algorithm to compress the file to be synchronized to obtain a compressed file, including the following steps:
[0170] S510, the synchronization tool creates a memory buffer and initializes a compression algorithm;
[0171] S520, the synchronization tool reads the to-be-synchronized file into memory in blocks, and compresses the data block of each to-be-synchronized file using a compression algorithm;
[0172] S530: The synchronization tool writes all compressed data blocks into a new file to obtain a compressed file.
[0173] A memory buffer is a preallocated, contiguous area of memory used to temporarily store data. The buffer size is determined based on the size of the data blocks you plan to process. For example, if you plan to process 1MB blocks of data at a time, you might allocate a larger buffer, such as 1.5MB, to allow for extra space to accommodate data expansion or other temporary needs. Memory buffers can be created using the memory allocation functions provided by your programming language.
[0174] Initializing a compression algorithm involves setting compression parameters and preparing the compression state. For example, using the DEFLATE algorithm, the initialization process includes: 1. Selecting a compression level (typically from 1 to 9, with 1 being the fastest but with the lowest compression ratio, and 9 being the slowest but with the highest compression ratio); 2. Setting the window size (which determines the range within which duplicate data is searched, typically 32KB); 3. Initializing the Huffman tree (used for dynamic Huffman encoding); and 4. Allocating and initializing internal data structures (such as hash tables, used to quickly find duplicate strings).
[0175] The file to be synchronized is read into memory in chunks and each chunk is compressed. Chunked reading is used to process large files that may exceed the available memory size, and also facilitates parallel processing and improves overall efficiency.
[0176] First, you need to determine an appropriate block size. Choosing a block size involves weighing several factors: memory usage, I / O efficiency, and compression efficiency. Typically, block sizes are between 64KB and several MB. For example, a block size of 1MB might be chosen. This size doesn't consume too much memory while still fully utilizing the read performance of modern hard drives.
[0177] The read process uses buffered I / O to improve efficiency. For example, if you have a 10MB file that needs to be compressed, using a 1MB block size, the read process might be as follows: 1. Open the file; 2. Loop 10 times, reading 1MB of data into a memory buffer each time; 3. Compress each 1MB block; 4. Process the final block, which may be less than 1MB.
[0178] After the to-be-synchronized files are read into the memory in blocks, a compression algorithm is used to compress the data blocks of each to-be-synchronized file. In this embodiment, the compression of each data block is independent.
[0179] All compressed data blocks are written to a new file. In this embodiment, the data compressed in the previous steps can be integrated into a complete compressed file. Specifically, first, a new file is created to store the compressed data. File creation usually uses the file creation API provided by the operating system. The writing process adopts a buffered write strategy to improve efficiency. Instead of writing to the disk immediately after compressing a data block, multiple compressed data blocks can be accumulated in memory and then written to the disk at one time to reduce the number of disk I / O operations and improve overall performance. For example, a 4MB write buffer can be set, and a write operation is performed only when the accumulated compressed data reaches or approaches 4MB.
[0180] During the writing process, necessary metadata must be added, including file size, compressed size, checksum, etc. Metadata is usually written at the beginning or end of the file. For example, a fixed-size header area (such as 512 bytes) can be reserved at the beginning of the file to store this metadata.
[0181] This implementation creates a memory buffer and initializes the compression algorithm, providing the necessary resources and environment for the compression process. The block-wise reading and compression strategy enables the solution to handle files of various sizes, facilitating parallel processing. The final writing step ensures the correct storage of compressed data and the integrity of its metadata. This not only significantly reduces file size and improves data transmission and storage efficiency, but also offers excellent scalability, making it suitable for a variety of file compression scenarios.
[0182] In one implementation of this embodiment, the synchronization tool encrypts the sensitive fields to obtain an encrypted data packet, including the following steps:
[0183] S610: For any sensitive field, the synchronization tool obtains the position of the sensitive field in the file to be synchronized and generates a random initialization vector;
[0184] S620: The synchronization tool uses a preset encryption algorithm to replace the sensitive field with a random initialization vector to obtain an encrypted sensitive field.
[0185] S630. The synchronization tool generates encryption metadata based on the random initialization vector and the position of the sensitive field in the file to be synchronized.
[0186] S640: The synchronization tool inserts the encrypted sensitive field into the corresponding position in the file to be synchronized to obtain an encrypted file;
[0187] S650: The synchronization tool packages the encrypted metadata and the encrypted file into a container file to obtain an encrypted data packet.
[0188] This embodiment first precisely locates the location of sensitive fields within the file. If the file to be synchronized is a structured data file (such as JSON or XML), a corresponding parser can be used to locate specific fields. If it is an unstructured text file, regular expressions or other pattern matching techniques can be used to locate sensitive fields. The location of sensitive fields includes the starting position (offset) and length of the field.
[0189] The random initialization vector (IV) is a fixed-length sequence of random bytes whose length depends on the encryption algorithm used. For example, for the AES encryption algorithm, the random initialization vector is usually 16 bytes (128 bits) long.
[0190] A preset encryption algorithm is used to replace the sensitive field with a random initialization vector to obtain an encrypted sensitive field. The preset encryption algorithm may be an AES encryption algorithm.
[0191] The encryption process first requires an encryption key. The encryption key is pre-set, and the key length depends on the selected encryption algorithm. For example, AES can use 128-bit, 192-bit, or 256-bit keys.
[0192] Taking AES-CBC mode as an example, the encryption process can be described as follows: 1. Pad sensitive field data to an integer multiple of the block size (AES uses a 16-byte block size); 2. Use the IV as the input for the first block; 3. Encrypt each data block: a. Perform an XOR operation on the current plaintext block and the previous ciphertext block (or IV); b. Process the XOR result through the AES encryption function; c. Use the output ciphertext block as the input for the next block; 4. Concatenate all encrypted blocks to form the final ciphertext. The XOR (exclusive OR) operation is a binary operation whose rule is: if the two compared bits are different, the result is 1; if they are the same, the result is 0.
[0193] For example, assuming the sensitive field is "secret123" (9 bytes), which needs to be padded to 16 bytes, the padded data might be: 73 65 63 72 65 74 31 32 33 07 07 07 07 07 07 07 (the last 7 bytes are all 0x07, indicating 7 bytes of padding).
[0194] Using the previously generated 16-byte IV and a hypothetical 256-bit key, the encryption process is as follows: 1. XOR the IV with the padded data; 2. Process the XOR result using the AES encryption function; 3. Output 16 bytes of ciphertext.
[0195] The resulting ciphertext might be (in hexadecimal): F8 3D A1 4F 9C 21 15 7B 89 5E 330A C2 D4 8F 1B. The 16-byte ciphertext is used to replace the original 9-byte sensitive field.
[0196] This embodiment converts sensitive fields into an encrypted form that cannot be directly read. By using a random IV and encryption algorithm, even if the same sensitive field is encrypted multiple times, different ciphertexts will be generated, greatly enhancing data security. The encrypted data cannot be deciphered without authorization, effectively protecting sensitive information. Furthermore, because encryption is performed on specific fields rather than the entire file, this method maintains the overall file structure, facilitating subsequent processing and decryption operations.
[0197] Generates encryption metadata based on the random initialization vector and the location of sensitive fields in the file to be synchronized to ensure subsequent correct decryption. The encryption metadata contains key information required for decryption but does not contain the actual sensitive data itself.
[0198] Encryption metadata may include random initialization vector (IV), location information of sensitive fields in the original file (starting offset and length), encryption algorithm and mode used (such as AES-CBC), length of encrypted data, checksum, etc.
[0199] Insert the encrypted sensitive fields into the corresponding position of the file to be synchronized to obtain the encrypted file. While keeping the overall structure of the file unchanged, only the sensitive part that needs to be encrypted can be replaced.
[0200] Using the previously obtained location information, replace the sensitive fields in the original file with the encrypted data. This can be achieved using slicing operations or byte array replacement methods.
[0201] For example, suppose the original sensitive field starts at byte 100 and is 20 bytes long, while the encrypted data is 32 bytes long. The replacement process could be: leave the first 100 bytes of the file unchanged, then insert the 32 bytes of encrypted data, and finally append the remaining content from byte 120 (i.e., 100 + 20) of the original file. This replaces the sensitive data while maintaining the integrity of the rest of the file.
[0202] Packaging encryption metadata and encrypted files into a container file ensures that the encrypted data can be properly transmitted and decrypted. First, select a container file format. Container file formats include ZIP, TAR, or a custom binary format. Using ZIP as an example, the packaging process can be divided into the following steps: 1) Create a new ZIP file; 2) Add the encrypted files to the ZIP file; 3) Add the encryption metadata as a separate file to the ZIP file.
[0203] This implementation first precisely locates sensitive fields and generates a random initialization vector, laying the foundation for subsequent encryption. An encryption algorithm is then used to convert sensitive data into ciphertext, ensuring data confidentiality. The generated encrypted metadata contains the key information required for decryption, facilitating subsequent decryption operations. Accurately inserting the encrypted data into the original file maintains the integrity of the file structure, allowing non-sensitive portions to remain accessible. Finally, by packaging the encrypted file and metadata, a self-contained, secure data package is created. This not only protects sensitive information but also maintains data availability and manageability.
[0204] In one implementation of this embodiment, the synchronization tool establishes an HTTPS connection with the cloud platform and uploads the encrypted data packet to the cloud platform in segments, including the following steps:
[0205] S710. The synchronization tool calculates a first file checksum of the encrypted data packet and generates upload metadata, wherein the upload metadata includes a first name, a file size, and a first file checksum, or the upload metadata includes a second name, a file size, and a first file checksum.
[0206] S720: The synchronization tool establishes an HTTPS connection with the cloud platform.
[0207] S730, the synchronization tool divides the encrypted data packet into multiple fixed-size data blocks, uploads each block to the cloud platform, and records the fixed-size data blocks that have been uploaded to the cloud platform;
[0208] S740: After the synchronization tool uploads all fixed-size data blocks to the cloud platform, the synchronization tool sends the upload metadata to the cloud platform.
[0209] First, a checksum algorithm is used to calculate the checksum of the entire encrypted data packet. The checksum algorithm can be SHA-256. This algorithm generates a 64-bit hexadecimal string representing the file's unique fingerprint. Next, upload metadata is generated. This metadata includes the primary or secondary name, the file size, and the primary file checksum. Upload metadata is formatted in JSON or XML.
[0210] To establish an HTTPS connection with the cloud platform, please refer to S440, which will not be repeated here in this application.
[0211] The upload process for splitting an encrypted data packet into multiple fixed-size chunks and uploading them one by one can be as follows: a. Establishing an HTTPS connection with the cloud platform. b. Multi-part upload: Split a 520KB file into two 260KB chunks. Upload the first chunk, and the server verifies its integrity. Upload the second chunk, and the server verifies it again. c. Sending upload metadata to the cloud platform.
[0212] The upload process uses a loop, reading and uploading data chunks one by one. Each chunk upload can be considered a separate HTTP POST request. In addition to the chunk itself, the request should also include metadata such as the chunk's sequence number and file identifier. This information enables the server to correctly assemble and manage received chunks.
[0213] During the upload process, the synchronization tool needs to keep track of the uploaded blocks. This can be achieved by maintaining a data structure (such as an array or bitmap) that marks the corresponding position of each successfully uploaded block. If the upload process is interrupted, the next restart can directly start from the unuploaded block, avoiding the repeated upload of the successful part.
[0214] After all fixed-size data blocks are uploaded, the upload metadata is sent to the cloud platform to provide the overall information of the file, so that the cloud platform can verify all received data blocks and correctly assemble them into a complete file.
[0215] First, the synchronization tool needs to confirm that all data blocks have been successfully uploaded. This can be achieved by checking the previously maintained data block upload status records. For example, if a Boolean array is used to record the upload status of each block, you should ensure that all elements in the array are true. If any blocks are found that have not been uploaded successfully, you should complete the upload of these blocks before proceeding to the next step. After confirming that all blocks have been uploaded, the synchronization tool encapsulates the previously generated upload metadata (including file name, file size, and file checksum) into an HTTP request and sends it to the cloud platform. This request usually uses the POST or PUT method, and the data format can be JSON or XML.
[0216] In addition to the metadata generated previously, this request can also include additional information such as the total number of chunks and the upload ID (a unique identifier assigned by the server when the upload starts). This helps the server-side perform file verification and assembly.
[0217] After receiving this request, the cloud platform performs a series of validation operations. First, it checks to see if all the data chunks it claims to have been received. Then, it combines all the chunks into a complete file, calculates the file's checksum, and compares it with the checksum provided in the upload metadata. If the checksums match, the file has been successfully uploaded in its entirety. If they do not, some or all of the chunks may need to be re-uploaded.
[0218] The synchronization tool waits for and correctly processes the server's response. Possible server responses include: upload success, verification failure requiring re-upload, partial block loss requiring retransmission, etc. Based on the response, the synchronization tool needs to take appropriate actions, such as reporting the upload success to the user or re-uploading the specified data block.
[0219] This implementation first calculates a file checksum and generates upload metadata, laying the foundation for subsequent file integrity verification. Establishing an HTTPS connection ensures data transmission security, preventing eavesdropping or tampering. The use of a block upload strategy significantly improves the reliability and efficiency of large file transfers, supports resumable uploads, and effectively addresses network instability. Finally, the upload metadata is sent as final confirmation, ensuring the integrity and consistency of the entire file.
[0220] Figure 5 A second flow chart of a method for managing digital electricity invoices based on an AI model provided in an embodiment of the present application is shown. Figure 5 As shown, the embodiment of the present application also provides a digital invoice management method based on an AI model, which is applied to a cloud platform, the cloud platform is connected to an electronic device, and the electronic device is connected to a physical printer. The method includes the following steps:
[0221] S801: In response to receiving a fixed-size data block transmitted by an electronic device, storing the fixed-size data block in a preset buffer until all fixed-size data blocks are received;
[0222] S802, reorganize all fixed-size data blocks in the buffer to obtain a reorganized file;
[0223] S803, calculating a second file checksum of the reorganized file;
[0224] S804: Receive and parse the uploaded metadata to obtain the first file verification code and the location of the sensitive field;
[0225] S805, comparing the second file checksum with the first file checksum;
[0226] S806. After the comparison is successful, the encrypted data packet is unpacked to obtain the encrypted metadata and the encrypted file, and the field to be decrypted is located according to the position of the sensitive field;
[0227] S807. Decrypt the field to be decrypted using a preset decryption algorithm and a preset security key, obtain the sensitive field corresponding to the field to be decrypted, and obtain the decrypted compressed file;
[0228] S808: Decompress the compressed file using a preset decompression algorithm to obtain an electronic business receipt;
[0229] S809: Parse the electronic business bill to obtain bill information, execute the billing process based on the bill information, and generate a digital invoice and callback instruction;
[0230] S810. Send the digital invoice and the callback instruction to the electronic device, so that the electronic device generates a printing instruction according to the callback instruction, and sends the printing instruction and the digital invoice to a physical printer, which is used to print the digital invoice according to the printing instruction.
[0231] In response to receiving fixed-size data blocks transmitted by an electronic device, the cloud platform will immediately store these data blocks in a pre-set buffer. This process will continue until all fixed-size data blocks have been received. The buffer is usually a temporarily allocated memory space used to temporarily store these data blocks. In a specific implementation, the cloud platform will maintain a data reception status, recording the number of data blocks received and the total number of data blocks to be received. Each time a data block is received, it is written to the corresponding position in the buffer and the reception status is updated. For example, assuming that each data block is 1MB in size and the total file size is 10MB, it will be divided into 10 data blocks for transmission. The cloud platform will allocate a 10MB buffer and write each 1MB data block to the corresponding position in the buffer (0-1MB, 1-2MB, ..., 9-10MB) when it is received.
[0232] All fixed-size data blocks in the buffer are reassembled to obtain a complete reassembled file. Specifically, data blocks stored in different locations in the buffer can be concatenated in the correct order to form a complete file. This task can be implemented using file pointers or memory operations. For example, if there are 10 1MB data blocks stored in the buffer, the reassembly process will start with the first block and write the contents of each block in sequence to a new file or memory space. This process requires precise control of the start and end positions of each block to ensure no data loss or duplication. For example, the first block would be written to the 0-1MB position of the new file, the second block to the 1-2MB position, and so on. If memory operations are used, data can be moved and concatenated directly in memory and then written to the file system in one go. Ultimately, the scattered data blocks are converted into a complete, usable file, ready for subsequent file integrity verification and decryption operations. This efficient reassembly process minimizes disk I / O and improves file processing speed.
[0233] A checksum algorithm is used to calculate the second file checksum of the reassembled file to ensure file integrity. In practice, the reassembled file contents are read block by block, each block input into the selected checksum algorithm, ultimately generating a checksum of fixed length. For example, when using the MD5 algorithm, a 128-bit (16-byte) checksum is always obtained regardless of file size. This ultimately generates a reliable file fingerprint that is subsequently compared with the first file checksum in the uploaded metadata to verify that the file has remained intact and has not been tampered with or damaged during the transmission and reassembly process. This method effectively detects potential errors during file transmission, improving overall system reliability and data security.
[0234] Receive and parse the uploaded metadata to obtain the first file verification code and the location of the sensitive fields. Specifically, the cloud platform receives the uploaded metadata from the electronic device via a network interface. After receiving the uploaded metadata, the platform uses a corresponding parser (such as a JSON parser) to extract the first file verification code and the location of the sensitive fields.
[0235] Comparing the second file checksum with the first file checksum can be achieved through a string comparison operation. The second file checksum is obtained by calculating the reconstructed file in step S803, while the first file checksum is parsed from the uploaded metadata in step S804. In specific implementation, it can be implemented using the string comparison function provided by the programming language. If the two checksums are not exactly the same, even if there is only one character difference, it will be considered a comparison failure. If the comparison is successful, it means that the file has not been tampered with or damaged during the transmission process and subsequent processing can be carried out. If the comparison fails, it means that an error may have occurred in the file during the transmission process and it needs to be retransmitted or other error handling measures need to be taken. Through this strict verification mechanism, the reliability and security of the system processing data can be greatly improved, and subsequent processing errors caused by file damage or tampering can be prevented.
[0236] After a successful checksum comparison, the encrypted data packet is unpacked to obtain the encrypted metadata and the encrypted file. The fields to be decrypted are then located based on the location of the sensitive fields. First, the encrypted data packet is parsed to separate the encrypted metadata and the encrypted file content. Then, based on the sensitive field location information previously obtained from the uploaded metadata, the data segments to be decrypted are precisely located within the encrypted file. For implementation, assume the structure of the encrypted data packet is: [encrypted metadata length (4 bytes)][encrypted metadata][encrypted file content]. First, the first 4 bytes are read to determine the length of the encrypted metadata. Then, the encrypted metadata of the corresponding length is read. The remaining portion represents the encrypted file content. For locating sensitive fields, assume there are two sensitive fields, located between bytes 100-200 and 300-400. The system marks these two areas within the encrypted file content as fields to be decrypted, thereby breaking down the complex encrypted data packet into individually processable components and accurately identifying the data segments to be decrypted.
[0237] Using a preset decryption algorithm and a preset security key, the fields to be decrypted are decrypted to obtain the sensitive fields corresponding to the fields to be decrypted, thereby obtaining the decrypted compressed file. First, the preset decryption algorithm (the reverse process of the encryption steps from steps S610 to S650) and the corresponding security key are loaded. Then, the previously located fields to be decrypted are decrypted one by one. In specific implementation, assume the use of the AES-256 algorithm and a pre-agreed 256-bit key. For each field to be decrypted, the corresponding encrypted data is read and decrypted using the AES decryption algorithm and key. For fields to be decrypted that are 100-200 bytes long, 100 bytes of data are read, decrypted using the AES-256 algorithm and key, and the decrypted data is then written back to its original location. This process is repeated for all marked fields to be decrypted. After decryption is complete, the sensitive information in the file is restored to its plaintext state. This embodiment can securely restore sensitive information in a file while leaving the rest of the file intact. By decrypting only the necessary portions, the security of sensitive information is ensured while improving processing efficiency.
[0238] Use the preset decompression algorithm to decompress the compressed file to obtain the electronic business bill. The decompression algorithm corresponds to the compression algorithm of S520. First, identify the format of the compressed file, and then select the corresponding decompression algorithm. The decompression process usually includes reading the header information of the compressed file, parsing the compressed data structure, and gradually decompressing each compressed block. In specific implementation, it is assumed that a compressed file in ZIP format is used. First, read the central directory of the ZIP file to obtain the file list and the compression information of each file. Then, decompression is performed on each compressed file item. This process is repeated for all items in the ZIP file until all files are decompressed. After the decompression is completed, one or more decompressed files are obtained, which are the complete electronic business bills.
[0239] The electronic business invoice is parsed to obtain invoice information. The invoicing process is then executed based on this invoice information, generating a digital invoice and a callback instruction. First, the appropriate parser is selected based on the electronic business invoice format (e.g., PDF, XML, etc.). The parsing process extracts key invoice information, such as the invoice header, amount, tax rate, and item details. Based on this extracted information, the invoicing process is then executed according to pre-defined invoicing rules. The parsed invoice information is then transmitted to the invoicing system. Based on pre-set rules and templates, the invoicing system fills the invoice information into the corresponding fields, generating an electronic invoice that complies with tax regulations. Simultaneously, a callback instruction is generated to control the subsequent printing process. For example, a 100-yuan restaurant receipt might be parsed to reveal a date of April 1, 2024, a product of "beef noodles," a quantity of two servings, and a unit price of 50 yuan. The invoicing system then generates an electronic invoice containing this information and creates a callback instruction containing printing parameters (e.g., paper size and number of copies).
[0240] The generated digital electronic invoice (abbreviated as "digital electronic invoice") and callback instructions are sent to the electronic device through a network transmission protocol (such as HTTPS). After receiving these data, the electronic device can first verify the integrity and authenticity of the data to ensure that it has not been tampered with. Then, the electronic device parses the callback instruction and generates a specific printing instruction based on the parameters in the instruction. This printing instruction contains detailed instructions on how to format and layout the content of the digital electronic invoice. For example, the instruction can specify the use of A4 paper, set a 1 cm page margin, use a 12-point Songti font, etc. After generating the print instruction, the electronic device packages the print instruction and digital electronic invoice data, and sends it to the physical printer through a local network or Bluetooth connection.
[0241] Finally, after receiving the print command and the digital invoice data, the physical printer first parses the print command to understand how to process the received data. The printer's control system then sets the printing parameters based on the command, such as selecting the appropriate paper size, adjusting the print head position, and setting the print resolution. The printer then converts the digital invoice content into a printable image or text format and prints it onto paper according to the specified layout. Once printing is complete, the printer can send a confirmation message to the electronic device confirming the successful printing of the invoice.
[0242] This embodiment automates the entire process, from electronic business bills to printed invoices. This effectively improves invoicing and printing efficiency, reduces human error, and ensures the accuracy and consistency of invoice information. Digital processing and transmission enhances data security and reduces the risk of information leakage. This not only simplifies a company's financial management process but also improves the user experience, making it easier to issue and obtain invoices.
[0243] In one implementation of this embodiment, the invoicing process includes the following steps:
[0244] S901. Call the preset tax interface, obtain the invoice code and invoice number based on the bill information, and generate an electronic signature;
[0245] S902. Fill in the invoice code, invoice number and electronic signature into the preset standard template to generate a digital invoice.
[0246] In the first step of the invoicing process, a pre-configured tax interface is called to obtain the invoice code and invoice number and generate an electronic signature. This process first involves a secure connection with the tax system. The connection uses the encrypted HTTPS protocol to ensure secure data transmission. Once the connection is established, the parsed invoice information is packaged into a specifically formatted request and sent to the tax interface. This request contains all the key information required for invoice issuance, such as seller and buyer information, a detailed description of the goods or services, and the transaction amount. Upon receiving the request, the tax system undergoes a series of verification and processing. It checks the legitimacy of the request, verifies the company's tax eligibility, and assigns a unique invoice code and invoice number based on the current invoice number segment.
[0247] After obtaining the invoice code and number, a predefined algorithm is used to generate an electronic signature to ensure the authenticity and integrity of the invoice. Specifically, the company's private key is used to encrypt the key invoice information (including the newly obtained invoice code and number) to generate a unique encrypted string, which serves as the electronic signature.
[0248] The captured invoice code, invoice number, and generated electronic signature are entered into a pre-set standard template to generate a complete digital electronic invoice (digital electronic invoice). Since different types of transactions may require different invoice templates, such as general VAT invoices and special VAT invoices, a template can be automatically selected based on previously parsed invoice information. These templates are typically pre-designed and meet the format and content requirements set by the State Administration of Taxation.
[0249] After selecting a template, fill in the various information in an orderly manner. First, enter basic information such as the invoice code and invoice number in the designated spaces. Next, enter the seller and buyer's detailed information, including name, taxpayer identification number, address, phone number, bank account, and account number. Next, fill in the corresponding columns with product or service information, including name, specification, unit, quantity, unit price, and amount. If there are multiple product items, the total amount will be automatically calculated and entered in the Total Amount column.
[0250] Finally, the previously generated electronic signature is added to the designated location on the invoice. This signature is usually presented as a QR code to facilitate subsequent verification. A standard QR code containing key invoice information is also generated for quick verification.
[0251] This implementation automatically captures the invoice code and number, generates an electronic signature, and populates this information into a standard template, significantly reducing human error and improving invoicing efficiency. Furthermore, the use of electronic signatures ensures the authenticity and immutability of invoices, enhancing their legal validity and credibility.
[0252] like Figure 6 As shown, the embodiment of the present application further provides a cloud platform, which is applied to the above-mentioned digital electricity invoice management method based on the AI model, including:
[0253] a communication module 10, configured to communicate with electronic devices; and
[0254] The data processing module 20 is connected to the communication module and is used to execute the above-mentioned digital electricity invoice management method based on the AI model from S801 to S810, and S901 to S902.
[0255] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0256] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0257] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0258] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0259] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0260] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0261] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0262] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0263] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A digital invoice management method based on an AI model, applied to an electronic device, wherein the electronic device is deployed with an AI model, a business platform, a virtual printer, and a synchronization tool, and the electronic device is connected to a physical printer, the method comprising: Responding to an invoicing signal from the business platform, obtaining an electronic business invoice generated by the business platform; Identify electronic business bills through AI models and obtain the domain type of the electronic business bills; Convert the electronic business bill into a file format using a virtual printer to obtain the converted electronic business bill, and store the converted electronic business bill in a corresponding folder; In response to an event trigger signal of the folder, the converted electronic business invoice is uploaded to the cloud platform through the synchronization tool, wherein the synchronization tool is used to monitor the folder, and the cloud platform is used to parse the converted electronic business invoice to obtain the invoice information, execute the invoicing process according to the invoice information, generate a digital invoice and a callback instruction, and send the digital invoice and the callback instruction to the synchronization tool, so that the synchronization tool generates a printing instruction according to the callback instruction, and the printing instruction is used to control the physical printer; Sending the printing instruction and the digital invoice to the physical printer, so that the physical printer prints the digital invoice according to the printing instruction; The steps to generate a folder include: Obtaining the first timestamp of the electronic business bill generation; If the domain type is the HIS domain, obtain the business code of the business platform, adopt the first preset naming rule, and generate the first name of the folder according to the first timestamp; Generate a first path of the folder based on a preset first basic path, a first timestamp and a business code; Based on the first name and the first path, a folder is created, and a HIS domain main folder and at least one HIS domain subfolder are created in the folder; Allocate a first permission for the HIS domain main folder and a second permission for at least one HIS domain subfolder, and record a second path of the HIS domain main folder and a second timestamp of creation of the HIS domain main folder; Create a first json file in the HIS domain main folder and generate a log, the first json file includes the second timestamp and domain type, and the log includes the second path, the first permission and the second permission; If the domain type is the ERP domain, obtain the business class code of the business platform, use the second preset naming rule, and generate a second name for the folder according to the first timestamp; Generate a third path of the folder based on the preset second basic path, the first timestamp and the business class code; Based on the second naming and the third path, a folder is created, and an ERP domain main folder and at least one ERP domain subfolder are created in the folder. The number of ERP domain subfolders is the same as that of HIS domain subfolders. Allocate a third permission for the ERP domain main folder and a fourth permission for at least one ERP domain subfolder, and record a fourth path for the ERP domain main folder and a third timestamp of creation of the HIS domain main folder; A second json file is created in the ERP domain main folder to generate a log. The second json file includes a third timestamp and domain type. The log includes a fourth path, a third permission, and a fourth permission.
2. The method according to claim 1, characterized in that The AI model is used to identify electronic business bills and obtain the field types of electronic business bills, including: Extract features of the electronic business bill to obtain text features and metadata features of the electronic business bill; Text features and metadata features are input into the AI model to obtain the domain type of the electronic business bill, where the domain type includes the HIS domain and the ERP domain.
3. The method according to claim 1, characterized in that Upload the converted electronic business invoices to the cloud platform through the synchronization tool, including: The synchronization tool uses the converted electronic business invoice as a file to be synchronized, and identifies the file type and file size of the file to be synchronized; The synchronization tool determines the corresponding compression algorithm according to the file type, and compresses the file to be synchronized using the compression algorithm to obtain a compressed file; The synchronization tool identifies the sensitive fields of the compressed file and encrypts the sensitive fields to obtain an encrypted data packet; The synchronization tool establishes an HTTPS connection with the cloud platform and uploads the encrypted data packets to the cloud platform in segments.
4. The method according to claim 3, characterized in that The synchronization tool uses a compression algorithm to compress the files to be synchronized, and obtains compressed files, including: The synchronization tool creates a memory buffer and initializes the compression algorithm; The synchronization tool reads the files to be synchronized into memory in blocks and compresses the data blocks of each file to be synchronized using a compression algorithm; The synchronization tool writes all compressed data blocks into a new file to obtain a compressed file.
5. The method according to claim 3, characterized in that The synchronization tool encrypts sensitive fields and obtains an encrypted data packet, including: For any sensitive field, the synchronization tool obtains the location of the sensitive field in the file to be synchronized and generates a random initialization vector; The synchronization tool uses a preset encryption algorithm to replace sensitive fields with random initialization vectors to obtain encrypted sensitive fields; The synchronization tool generates encrypted metadata based on the random initialization vector and the position of sensitive fields in the files to be synchronized; The synchronization tool inserts the encrypted sensitive fields into the corresponding locations of the files to be synchronized to obtain the encrypted files; The synchronization tool packages the encryption metadata and the encrypted files into a container file to obtain an encrypted data package.
6. The method according to claim 5, characterized in that The synchronization tool establishes an HTTPS connection with the cloud platform and uploads encrypted data packets to the cloud platform in segments, including: The synchronization tool calculates a first file checksum of the encrypted data packet and generates upload metadata, wherein the upload metadata includes the first name, the file size, and the first file checksum, or the upload metadata includes the second name, the file size, and the first file checksum; The synchronization tool establishes an HTTPS connection with the cloud platform; The synchronization tool divides the encrypted data packet into multiple fixed-size data blocks, uploads them to the cloud platform one by one, and records the fixed-size data blocks that have been uploaded to the cloud platform; After the synchronization tool uploads all fixed-size data blocks to the cloud platform, the synchronization tool sends the upload metadata to the cloud platform.
7. A digital electricity invoice management method based on AI model, characterized by: The method is applied to the cloud platform according to any one of claims 1 to 6, wherein the cloud platform is connected to an electronic device, and the electronic device is connected to a physical printer, and the method comprises: In response to receiving a fixed-size data block transmitted by the electronic device, storing the fixed-size data block in a preset buffer until all fixed-size data blocks are received; Reassemble all fixed-size data blocks in the buffer to obtain a reassembled file; Calculating a second file checksum of the reassembled file; Receive and parse the uploaded metadata to obtain the first file checksum and the location of sensitive fields; comparing the second file checksum with the first file checksum; After the comparison is successful, the encrypted data packet is unpacked to obtain the encrypted metadata and the encrypted file, and the field to be decrypted is located based on the location of the sensitive field; Decrypt the field to be decrypted using a preset decryption algorithm and a preset security key, obtain the sensitive field corresponding to the field to be decrypted, and obtain the decrypted compressed file; Decompress the compressed file using a preset decompression algorithm to obtain an electronic business receipt; Parse electronic business bills to obtain bill information, execute the billing process based on the bill information, and generate digital invoices and callback instructions; The digital invoice and callback instruction are sent to the electronic device so that the electronic device generates a printing instruction according to the callback instruction, and sends the printing instruction and the digital invoice to the physical printer, which is used to print the digital invoice according to the printing instruction.
8. The method according to claim 7, characterized in that The invoicing process includes: Call the preset tax interface, obtain the invoice code and invoice number based on the bill information, and generate an electronic signature; Fill in the invoice code, invoice number and electronic signature into the preset standard template to generate a digital invoice.
9. A cloud platform, characterized in that: The digital electricity invoice management method based on the AI model applied to any one of claims 1-6, and the digital electricity invoice management method based on the AI model applied to any one of claims 7-8, comprising: a communication module for communicating with an electronic device; and The data processing module is connected to the communication module and is used to execute the digital electricity invoice management method based on the AI model according to any one of claims 7 to 8.
Citation Information
Patent Citations
Financial management electronic bill classified storage method and system
CN119399776A
System and method for parsing visual information to extract data elements from randomly formatted digital documents
US20210374190A1