Intelligent tax refund declaration method and device for cross-border trade, computer equipment and medium
Through adaptive analytical matrix and knowledge graph technology, we can solve the problem of data heterogeneity in cross-border trade, realize the automation and accuracy of cross-border trade intelligent tax refund declaration, and enhance the competitiveness of enterprises in international trade.
Patent Information
- Application Number
- CN202510869007.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-17
AI Technical Summary
Enterprises in cross-border trade face the problem of data heterogeneity in the process of export tax rebate declaration, which leads to difficulties in data integration, high manpower consumption, inaccurate and error-prone declarations, affecting timeliness and compliance.
Adaptive parsing matrix is used to extract tax refund data features, build a knowledge graph mapping relationship library, and combine optical character recognition, natural language processing and risk assessment models to automatically fill in declaration templates and perform tax refund operations.
It improves the efficiency and accuracy of tax refund declarations, reduces human errors, ensures data consistency and compliance, and enhances the competitiveness of enterprises in international trade.
Smart Images

Figure CN120807181A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent tax refund declaration of cross-border trade, and in particular to an intelligent tax refund declaration method and device for cross-border trade, a computer device and a medium. BACKGROUND
[0002] Under the background of globalization trade, export tax refund, as an important policy to promote the international competitiveness of enterprises, has attracted more and more attention from enterprises. However, in the actual process of export tax refund declaration, enterprises face two core pain points, one of which is the data heterogeneity problem. This problem is caused by the use of various information systems by enterprises in various links, such as ERP systems (such as SAP), logistics platforms (such as Flexport), and electronic invoice systems (such as value-added tax anti-fraud tax control). The data format and structure of these systems are different, which makes it difficult to integrate data.
[0003] Traditional processing methods perform well in processing structured data, but often appear to be inadequate when facing unstructured data, which makes enterprises have to spend a lot of manpower and time to manually compare and arrange data when declaring export tax refund. In addition, this process is prone to human error, affecting the accuracy and timeliness of the declaration, which may cause the enterprise to face the risk of delayed tax refund or policy compliance. SUMMARY
[0004] Therefore, it is necessary to propose an intelligent tax refund declaration method, device, computer device and medium for cross-border trade to solve the existing problems of intelligent tax refund declaration of cross-border trade.
[0005] An intelligent tax refund declaration method for cross-border trade, the method comprising:
[0006] obtaining a specified tax refund item and a plurality of tax refund data related to the specified tax refund item;
[0007] extracting features of each of the tax refund data through a preset adaptive analysis matrix to obtain a plurality of commodity features;
[0008] based on each commodity feature, obtaining and constructing a mapping relationship database of commodity features and tariff codes through a preset knowledge graph;
[0009] extracting the declaration content in the mapping relationship database and filling the declaration content into a preset declaration template to obtain a tax refund declaration form;
[0010] based on the tax refund declaration form, performing a tax refund operation.
[0011] Further, before the step of obtaining and constructing a mapping relationship database of commodity features and tariff codes through a preset knowledge graph based on each commodity feature, the method further comprises:
[0012] Recognize the data information of each commodity to be traded through preset optical characters;
[0013] Identify data information through a preset natural language processing model to obtain various entities and the relationships between them;
[0014] Each entity is used as a node, and the relationship between each entity is stored in a preset graph database to obtain the preset knowledge graph.
[0015] Furthermore, before the step of extracting the declaration content from the mapping relationship library and filling the declaration content into a preset declaration template to obtain a tax refund declaration form, the step further includes:
[0016] Performing semantic analysis on the designated tax refund items to extract the exporting country category;
[0017] Based on the exporting country category, the preset declaration template is obtained from a preset template database.
[0018] Furthermore, before the step of performing the tax refund operation based on the tax refund declaration form, the method further includes:
[0019] Constructing a multidimensional evaluation matrix for the tax refund declaration form based on the XGBoost algorithm;
[0020] Inputting the multidimensional assessment matrix into a preset risk assessment model to obtain a corresponding risk level;
[0021] If the risk level is greater than the preset risk level, the tax refund declaration form will be sent to a designated terminal for review.
[0022] Furthermore, before the step of inputting the multidimensional assessment matrix into a preset risk assessment model to obtain the corresponding risk level, the step further includes:
[0023] Obtain a specified number of customs audit cases from a preset case database, and divide the specified number of customs audit cases into a training data set and a validation data set according to a preset ratio;
[0024] Inputting the cases in the training data set into the initial risk assessment model for training to obtain a temporary risk assessment model;
[0025] Validating the temporary risk assessment model using cases in the validation dataset;
[0026] If the verification result is passed, the temporary risk assessment model is identified as the preset risk assessment model.
[0027] Further, the step of extracting features of each of the tax refund data through a preset adaptive parsing matrix to obtain a plurality of commodity features comprises:
[0028] Obtaining a data type of each of the tax refund data;
[0029] Based on the data type of each of the tax refund data, obtaining a corresponding parsing matrix;
[0030] Inputting the parsing matrix into a preset large language model to obtain a parsing model of the corresponding data type;
[0031] Parsing the corresponding tax refund data through the parsing model of each data type to obtain a plurality of commodity features.
[0032] Further, the step of obtaining and constructing a mapping relationship database of commodity features and tariff codes based on each commodity feature through a preset knowledge graph further comprises:
[0033] Monitoring in real time whether the tax refund policy has changed through a preset crawler;
[0034] If a change has occurred, obtaining a first target text after the change and a second target text before the change;
[0035] Calculating the similarity between the first target text and the second target text;
[0036] Determining whether the similarity is greater than a threshold value;
[0037] If the similarity is greater than the threshold value, constructing a GNN network to simulate a conduction effect to obtain a predicted entity affected by the prediction, and updating the relationship between each entity in the initial knowledge graph to obtain an updated preset knowledge graph.
[0038] An intelligent tax refund declaration device for cross-border trade, the device comprising:
[0039] An acquisition module for acquiring a specified tax refund item and a plurality of tax refund data related to the specified tax refund item;
[0040] An extraction module for extracting features of each of the tax refund data through a preset adaptive parsing matrix to obtain a plurality of commodity features;
[0041] A construction module for obtaining and constructing a mapping relationship database of commodity features and tariff codes based on each commodity feature through a preset knowledge graph;
[0042] A filling module for extracting declaration content in the mapping relationship database and filling the declaration content into a preset declaration template to obtain a tax refund declaration form;
[0043] A tax refund module is configured to perform a tax refund operation based on the tax refund declaration form.
[0044] A computer device includes a memory and a processor, the memory storing a computer program, the computer program being executed by the processor to cause the processor to perform the following steps:
[0045] Obtaining a specified tax refund item and a plurality of tax refund data related to the specified tax refund item;
[0046] Extracting features of each of the tax refund data through a preset adaptive analysis matrix to obtain a plurality of commodity features;
[0047] Based on each commodity feature, a mapping relationship database of commodity features and tariff codes is obtained and constructed through a preset knowledge graph;
[0048] Extracting the declaration content in the mapping relationship database and filling the declaration content into a preset declaration template to obtain a tax refund declaration form;
[0049] Based on the tax refund declaration form, a tax refund operation is performed.
[0050] A computer-readable storage medium stores a computer program, the computer program being executed by a processor to cause the processor to perform the following steps:
[0051] Obtaining a specified tax refund item and a plurality of tax refund data related to the specified tax refund item;
[0052] Extracting features of each of the tax refund data through a preset adaptive analysis matrix to obtain a plurality of commodity features;
[0053] Based on each commodity feature, a mapping relationship database of commodity features and tariff codes is obtained and constructed through a preset knowledge graph;
[0054] Extracting the declaration content in the mapping relationship database and filling the declaration content into a preset declaration template to obtain a tax refund declaration form;
[0055] Based on the tax refund declaration form, a tax refund operation is performed.
[0056] The beneficial effects of the present application are: using an adaptive analysis matrix to extract features of tax refund data, which can flexibly cope with different data types, optimize the structure and content of information, and provide high-quality input for subsequent analysis, through the knowledge graph constructed based on commodity features, the required declaration information is quickly obtained, reducing the complexity and delay in the traditional tax refund declaration process, thereby improving the declaration efficiency, ensuring the consistency and accuracy of the data, effectively reducing the compliance risk caused by human errors, enabling enterprises to perform tax refund operations more quickly and accurately, and enhancing their competitiveness and flexibility in international trade. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0058] in:
[0059] Figure 1 This is a diagram of an application environment of an intelligent tax refund declaration method for cross-border trade in one embodiment;
[0060] Figure 2 A flowchart of an intelligent tax refund declaration method for cross-border trade in one embodiment;
[0061] Figure 3 This is a structural block diagram of an intelligent tax refund declaration device for cross-border trade in one embodiment;
[0062] Figure 4 FIG. 1 is a structural block diagram of a computer device in one embodiment. DETAILED DESCRIPTION
[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0064] Figure 1 This is a diagram of the application environment for intelligent tax refund declaration for cross-border trade in one embodiment. Figure 1 This intelligent tax refund declaration method for cross-border trade is applied to an intelligent tax refund declaration system for cross-border trade. This intelligent tax refund declaration system for cross-border trade includes a terminal 110 and a server 120. Terminal 110 and server 120 are connected via a network. Terminal 110 can be a desktop terminal or a mobile terminal. The mobile terminal can be at least one of a mobile phone, tablet computer, or laptop computer. Server 120 can be implemented as a standalone server or a server cluster consisting of multiple servers. Terminal 110 is used to obtain tax refund items, and server 120 is used to perform tax refund operations.
[0065] like Figure 2As shown, in one embodiment, an intelligent tax refund declaration method for cross-border trade is provided. The method can be applied to both terminals and servers, and the present embodiment is exemplified by application to servers. The intelligent tax refund declaration method for cross-border trade specifically includes the following steps:
[0066] S1: Obtain a specified tax refund item and a plurality of tax refund data related to the specified tax refund item;
[0067] S2: Extract features of each of the tax refund data through a preset adaptive analysis matrix to obtain a plurality of commodity features;
[0068] S3: Based on each commodity feature, obtain and construct a mapping relationship database of commodity features and tariff codes through a preset knowledge graph;
[0069] S4: Extract the declaration content in the mapping relationship database and fill the declaration content into a preset declaration template to obtain a tax refund declaration form;
[0070] S5: Perform a tax refund operation based on the tax refund declaration form.
[0071] As described in step S1 above, obtain a specified tax refund item and a plurality of tax refund data related to the specified tax refund item. First, the specific items for which tax refund is needed must be clarified. These items can include different types of goods, services or related expenses. Enterprises usually rely on systems such as ERP systems or financial software to identify these tax refund items and extract a plurality of tax refund data related thereto. These data can include sales invoices, transportation documents, purchase records and other relevant financial documents. For cross-border trade enterprises, the amount and variety of relevant data are huge and diversified, involving different regions and the applicability of different laws and regulations. Therefore, comprehensive data needs to be obtained.
[0072] As described in step S2 above, features of each of the tax refund data are extracted through a preset adaptive analysis matrix to obtain a plurality of commodity features. After successfully obtaining the tax refund data, the second step is to analyze these data in depth through an adaptive analysis matrix to extract valuable commodity features. The adaptive analysis matrix is a mathematical tool that can be used to structure and analyze tax refund data. Specifically, key information in the data such as product name, quantity, unit price, etc. needs to be identified and converted into usable feature information. In this process, a certain degree of intelligence and flexibility should be possessed to automatically adjust the feature extraction algorithm for different data types, ensuring the accuracy and real-time performance of the extraction results. Through sufficient analysis of the tax refund data, enterprises can obtain detailed feature data about each product, which provides a basis for subsequent steps and ensures that the establishment of subsequent mapping relationships accurately reflects the actual situation.
[0073] As described in step S3 above, based on each commodity feature, a mapping relationship library of commodity features and tariff codes is obtained and constructed through a preset knowledge graph. According to the extracted commodity features, a set of mapping relationship library of commodity features and tariff codes is constructed through a preset knowledge graph. The knowledge graph is a semantic network constructed by nodes and edges, where the nodes represent specific commodity features, and the edges represent the relationship between these features and the tariff codes. By integrating industry standards and tax regulations, each commodity feature can be assigned a corresponding tariff code. This process involves data matching and logical reasoning, so that the establishment of the mapping relationship not only depends on simple database queries, but also emphasizes intelligent data analysis.
[0074] As described in step S4 above, the declaration content in the mapping relationship library is extracted and filled into the preset declaration template to obtain the tax refund declaration form. After successfully obtaining and constructing the mapping relationship library, the required declaration content is extracted, and according to the requirements of the enterprise, the relevant commodity features and tariff codes are automatically extracted from the mapping relationship library, and the preset tax refund declaration template is filled according to these information. The template generally contains necessary declaration fields, such as tax refund items, commodity details, quantity, tax refund amount, etc. Under the premise of ensuring the accuracy of the filled content, the complete tax refund declaration form is quickly generated by using automatic technology. In this process, the compliance and completeness of the reported information also need to be ensured to prevent the failure of declaration due to missing or incorrect information. Through automatic reporting, the enterprise not only saves a lot of labor cost, but also improves the efficiency of declaration and shortens the period of applying for tax refund.
[0075] As described in step S5 above, the tax refund operation is performed based on the tax refund declaration form. The enterprise will submit the declaration form to the relevant tax department or agency for formal tax refund application. The process of tax refund operation not only covers file submission, but also may involve further data verification and audit. At this time, the status of tax refund application can be tracked in real time, and corresponding support can be provided when necessary to deal with possible inquiries or audits. By optimizing the process of tax refund operation, the enterprise can improve the success rate and timeliness of the entire declaration. In addition, during this process, the enterprise should maintain good communication with the tax authorities to solve any potential problems and ensure timely receipt of the due tax refund amount. Finally, the intelligent tax refund declaration method integrating the above steps helps the enterprise effectively manage the tax refund risk in cross-border trade and improve competitiveness.
[0076] In one embodiment, before the step S3 of obtaining and constructing the mapping relationship library of commodity features and tariff codes based on each commodity feature through a preset knowledge graph, the method further comprises:
[0077] S201: obtaining data information of each commodity to be traded through a preset optical character recognition;
[0078] S202: Identify the data information by a preset natural language processing model, to obtain each entity and the relationship between each entity;
[0079] S203: Take each entity as a node, and store the relationship between each entity in a preset graph database, to obtain the preset knowledge graph.
[0080] As described in the above steps S201-S203, the data information of the trade commodity is extracted using optical character recognition (OCR) technology. OCR is a technology that converts text in images into editable and searchable text format, widely used in scenarios such as scanning documents and identifying product labels. For cross-border trade, enterprises need to extract important information such as product name, quantity, price, invoice number, etc. from different documents, which may include paper invoices, transportation documents and other related certificates. Through the preset OCR system, the information in these paper documents can be automatically converted into machine-readable format, thereby improving the efficiency of data processing. The OCR system first preprocesses the input image, including denoising, rotation correction, etc. to improve the clarity of the text. Then, the algorithm analyzes the text in the image and converts it into digital information. Finally, the text information generated by the system is verified and proofread to ensure accuracy.
[0081] After completing the OCR processing of the data, the extracted data information is further analyzed using a natural language processing (NLP) model. In this step, the NLP model performs word segmentation, entity recognition and relationship extraction on the obtained text data, thereby identifying key information and entities in the text. For example, enterprise name, product name, transaction method, amount, etc. can be considered as independent entities. At the same time, the NLP system can identify the relationship between these entities, such as a certain enterprise buying a certain product, the transaction amount, etc. The NLP model usually includes multiple processing components, such as named entity recognition, syntax analysis and semantic analysis, which work together to accurately capture the meaning of the text content. Through training the model, the system can continuously improve the recognition ability of industry-specific terminology, and further extract more valuable context relationships from the data information. The importance of this process lies in that it can refine structured content from raw data information, providing the required nodes and edges for building a knowledge graph. In addition, based on the extracted entities and their relationships, enterprises can further conduct data analysis to identify the circulation mode of goods and other potential business insights.
[0082] The identified entities and their relationships are structured and stored, and a preset knowledge graph is constructed. The knowledge graph is a knowledge base represented in the form of a graph, with nodes representing entities (such as goods, enterprises, transactions, etc.) and edges representing relationships between different entities (such as buying and selling, production, supply chain relationships, etc.). Storing entities and relationships in a graph database enables efficient data querying and relationship reasoning, facilitating subsequent data analysis and decision support. Graph databases are typically stored and queried in graph structures, and the construction rules of the graph need to be determined, including node attributes, edge weights, and other information. Each node will contain relevant metadata, such as the name, price, and classification of a product, and the edges will specify the type and direction of the relationship. At the same time, enterprises can use the features of the preset graph database to quickly obtain all nodes and relationships related to a specific product, supporting real-time data analysis and modeling. This process lays the foundation for constructing a mapping relationship library between product features and tariff codes, making subsequent data processing more efficient and accurate, and also providing support for intelligent decision-making in cross-border trade. In this way, enterprises can better understand the overall picture of their products in the process of seeking tax refunds, thereby improving compliance and efficiency in complex tax environments.
[0083] In one embodiment, before the step S4 of extracting the declaration content in the mapping relationship library and filling the declaration content into the preset declaration template to obtain a tax refund declaration form, the method further comprises:
[0084] S301: performing semantic analysis on the specified tax refund item to extract the export country category;
[0085] S302: based on the export country category, obtaining the preset declaration template from a preset template database.
[0086] As described in steps S301-S302 above, semantic analysis is first performed on the designated tax refund item to extract category information related to the exporting country. Semantic analysis is a key technology in natural language processing (NLP). It not only focuses on the literal meaning of words, but also examines the relationships between words and the overall meaning of sentences. In cross-border trade, tax refund items often need to be classified according to the laws and regulations of different exporting countries to ensure compliance and accuracy of declarations. Therefore, extracting the exporting country category is particularly important. Specifically, a preset semantic analysis model is used to process the tax refund item description to identify relevant keywords, phrases, or features. By performing operations such as word segmentation and named entity recognition on the input text, the semantic analysis model can automatically identify the name or related identifiers of the exporting country. In addition to extracting information on a single exporting country, it can also handle complex situations involving multi-country transactions and summarize appropriate categories. For example, if a tax refund item involves transactions with multiple exporting countries, contextual analysis can be used to reasonably categorize and generate the corresponding exporting country category. This analysis not only provides the necessary information for subsequent steps but also ensures the targeted and efficient subsequent data processing. After extracting the exporting country category, the corresponding declaration template is retrieved from a pre-set template database based on the identified exporting country category. Tax refund policies, declaration formats, and requirements may vary from country to country. To comply with relevant laws and regulations, companies must select a format that matches the exporting country. The pre-set template database contains tax refund declaration form templates for multiple countries and regions, categorized by exporting country category. This pre-generated template database automatically matches and retrieves the correct template based on the specific rules of the exporting country. After selecting a template, the template's field settings are further checked to ensure that the required declaration information is accurately entered. The template may contain placeholders for information such as product name, quantity, amount, tariff code, and exporting country information, which will be filled in later steps. This process enables companies to prepare tax refund applications quickly and compliantly, reducing the risk of delays or errors. This automated process also effectively reduces the burden of manual operations and improves data processing efficiency.
[0087] In one embodiment, before step S5 of performing the tax refund operation based on the tax refund declaration form, the method further includes:
[0088] S401: Constructing a multidimensional evaluation matrix for the tax refund declaration form based on the XGBoost algorithm;
[0089] S402: Inputting the multi-dimensional assessment matrix into a preset risk assessment model to obtain a corresponding risk level;
[0090] S403: If the risk level is greater than the preset risk level, the tax return form is sent to a designated terminal for review.
[0091] In one embodiment, before the step S402 of inputting the multi-dimensional evaluation matrix into the preset risk assessment model to obtain the corresponding risk level, the method further comprises:
[0092] S4011: Obtain a specified number of customs inspection cases from a preset case database, and divide the specified number of customs inspection cases into a training data set and a validation data set according to a preset ratio;
[0093] S4012: Input the cases in the training data set into an initial risk assessment model for training to obtain a temporary risk assessment model;
[0094] S4013: Verify the temporary risk assessment model using the cases in the validation data set;
[0095] S4014: If the verification result is passed, the temporary risk assessment model is identified as the preset risk assessment model.
[0096] As described in steps S4011-S4014 above, a certain number of relevant cases will be extracted from the preset customs inspection case database first. These cases can be past inspection records, audit results or other review results, which provide valuable data to help the training and verification of the model. Preferably, the customs inspection cases in recent years. The extracted cases need to be divided according to a preset ratio, into a training data set and a validation data set. Usually, in order to ensure the effectiveness of model training and verification, the proportion of the training data set should be higher, for example, 70% to 80%, and the validation data set accounts for 20% to 30%. The purpose of this division method is to avoid overfitting of the model, so that the performance of the model on unseen data can be effectively evaluated. Throughout the process, the enterprise also needs to ensure that the selected cases are representative, and can cover a variety of data features and different risk situations, thereby providing a solid foundation for subsequent risk assessment model construction.
[0097] After dividing the training data set and the validation data set, the cases in the training data set are input into the initial risk assessment model for training. The initial risk assessment model can be a basic algorithm model such as logistic regression, decision tree, random forest, etc., or a more complex deep learning model, depending on the needs and available data of the enterprise. The training process of the model learns the input case data through algorithms, including identifying features, evaluating the contribution of each feature to the risk of inspection, etc. During the training process, the model continuously adjusts its internal parameters to minimize the defined loss function. Finally, after training is completed, the model outputs a preliminary evaluation result, which is referred to as a temporary risk assessment model. This model is optimized on a specific training set, but may not have been validated, so the key at this stage is to ensure that the model can effectively identify and assess customs inspection risks while preventing over-idealization on specific data sets, which affects the generalization ability of the model. After completing the training of the temporary risk assessment model, the performance of the temporary risk assessment model is evaluated using the divided validation data set. The validation data set is data that has not been seen during the model training process. The significance of using these data to evaluate the model lies in judging the generalization ability and actual application effect of the model. Specifically, the validation process includes inputting the cases in the validation data set into the temporary risk assessment model, comparing the evaluation results output by the model with the actual inspection results, and calculating multiple performance indicators such as accuracy, recall rate, F1 value, etc. of the model. These indicators will help the enterprise understand the performance of the model on new data and evaluate its prediction ability and effectiveness for customs inspection risks. If the model performs well on the validation data set and meets the preset performance standards, the enterprise will have confidence that the model has the ability to be applied in practice; if it fails to pass the validation, the model may need to be adjusted, retrained or the diversity of the data set increased to improve the prediction accuracy of the model. After validation, whether the temporary risk assessment model is converted into a preset risk assessment model is determined according to the validation result. If the performance of the model in the validation phase meets the preset standards and shows good accuracy and stability, the model is recognized as a mature and reliable risk assessment tool and is defined as a preset risk assessment model. Thus, only models that have been validated and have excellent performance can be used in actual applications. In the case where the validation result is passed, the enterprise will put the model into use and apply it to real-time risk assessment to help identify potential customs inspection risks and provide valuable data support at the decision-making level.
[0098] In one embodiment, the step S2 of extracting features of each of the tax refund data by the preset adaptive analysis matrix comprises:
[0099] S211: obtaining the data type of each of the tax refund data;
[0100] S212: Obtain a corresponding analysis matrix based on the data type of each of the tax refund data;
[0101] S213: Input the analysis matrix into a preset large language model to obtain an analysis model corresponding to the data type;
[0102] S214: Analyze the corresponding tax refund data through the analysis model of each data type to obtain a plurality of commodity characteristics.
[0103] As described in steps S211-S214 above, it is necessary to identify and obtain the data types of each tax refund data. Tax refund data can contain various formats, such as text files, tables, images, or other electronic documents, and each data type may require different methods and tools for processing and parsing. Through data analysis and feature exploration, metadata extraction techniques can be used to identify data types. This process is achieved by checking file extensions, data structure, content, and format. For example, PDF documents, CSV files, image files, and XML files can be identified separately. For each format, further analysis of its content can be performed to ensure that the organization form and type of information can be accurately obtained. This operation not only involves the use of program algorithms, but also may include API calls to different data sources, database queries, etc. After identifying the data type, a solid foundation is laid for subsequent selection of appropriate parsing algorithms and tools. Different types of data often require different methods to extract useful features. The parsing matrix is a set of rules and algorithms defined for a specific data type, designed to guide how to extract effective information and features from the data. Specifically, a set of parsing matrices can be pre-set, each of which is designed for different data types (such as text, image, structured data, etc.), detailing the specific steps and rules of feature extraction. For example, for text data, the parsing matrix may define tokenization, stop word removal, keyword extraction, etc.; for table data, the parsing matrix may describe column recognition, value association analysis, etc. The obtained parsing matrix is input into the pre-set large language model to generate a parsing model corresponding to the specific data type. Large language models (such as GPT, BERT, etc.) have powerful natural language understanding and generation capabilities, and can provide deep contextual understanding and semantic analysis when parsing complex text and other forms of data. The process of inputting the parsing matrix into the large language model usually involves multiple levels. The model first parses the explicit rules and features, and then identifies and strengthens them through its own learning ability. This is an iterative process, where the model will use pre-trained knowledge to understand the content in the parsing matrix while adapting to specific parsing tasks. In this way, enterprises can not only take advantage of the powerful capabilities of existing large language models, but also tailor them to specific data types to generate customized parsing models. After obtaining parsing models for different data types, these models are used to parse the corresponding tax refund data to extract multiple product features. Specifically, the downloaded tax refund data is input into the corresponding parsing model one by one. For example, for text data, the parsing model may perform natural language processing to extract product name, quantity, price, and other key features; for structured data, the model will parse the relevant field content according to the established rules; and for image data, deep learning techniques can be used to extract the recorded product information from the image.After this step, the system will be able to generate a set of product features that will serve as the basis for subsequent tax refund declarations, risk assessments, and decision support. This extraction process not only improves data processing efficiency and consistency, but also ensures the integrity and accuracy of the information. In addition, through this intelligent feature extraction method, enterprises can quickly respond to market changes and compliance requirements, further enhancing their competitiveness in cross-border trade.
[0104] In one embodiment, before the step S3 of obtaining and constructing a mapping relationship library between product features and tariff codes based on individual product features through a preset knowledge graph, the method further comprises:
[0105] S221: Real-time monitoring of changes in tax refund policies through a preset crawler;
[0106] S222: If changes have occurred, obtaining the first target text after the changes and the second target text before the changes;
[0107] S223: Calculating the similarity between the first target text and the second target text;
[0108] S224: Determining whether the similarity is greater than a threshold value;
[0109] S225: If the similarity is greater than the threshold value, constructing a GNN network to simulate the conduction effect, obtaining a predicted entity affected by the prediction, and updating the relationship between each entity in the initial knowledge graph to obtain an updated preset knowledge graph.
[0110] As described in steps S221-S225 above, a preset crawler program is deployed to monitor changes in historical tax refund policies in real time. A crawler is an automated script that periodically accesses specific websites or databases to extract and update data. In cross-border trade, tax refund policies may change, but it is uncertain whether the changes will affect automated tax refund operations. Therefore, the preset crawler will be configured to periodically access relevant data sources to check for updates to tax refund policies. These data sources may include customs official websites, trade-related regulatory websites, and industry dynamic update platforms. When changes in data are identified, the operation of obtaining the latest data and the previously archived data will be triggered. That is, the first target text and the second target text are obtained, with the first target text referring to the latest tax refund policy and the second target text representing the original data before the change. The data structures of these two texts may be similar, but their contents may differ, such as new clauses, added product tariffs, or deleted regulations. After obtaining both, the enterprise will be able to compare and analyze the changes and understand the impact of these changes on the current knowledge graph and the mapping relationship between product features and tariff codes.
[0111] The similarity between the two texts is calculated. Similarity calculation can use various algorithms, such as cosine similarity, Jaccard index, edit distance, etc. These algorithms analyze the same or similar words, phrases, and sentence structures in the two texts to quantify their similarity. In this way, the enterprise can accurately locate which parts have changed significantly, and then understand the potential impact of these changes on business operations, tax policies or compliance requirements. A higher similarity value may indicate limited changes, and the enterprise can operate based on the original strategy; while a lower similarity value may indicate the need to re-evaluate more far-reaching policy changes. After calculating the similarity, the enterprise needs to determine whether the similarity value is higher than the preset threshold. Setting the threshold is a key strategy that usually needs to be determined based on industry standards, the specific needs of the enterprise and historical data analysis. If the similarity is greater than the preset threshold, it means that the change is not significant, and the enterprise can choose to maintain the existing strategy and continue to use the previous knowledge graph and mapping relationship; if the similarity is less than the threshold, it indicates that the tax refund policy has changed significantly, which needs to be analyzed and evaluated in depth.
[0112] After confirming that the similarity is greater than the threshold, a graph neural network (GNN) is constructed to simulate and analyze the transmission effect of the change in the tax refund policy and predict its impact on other entities in the established knowledge graph. GNN is a deep learning algorithm suitable for processing graph-structured data, which can effectively capture the relationships between nodes and propagate and update information through these relationships. When constructing the GNN network, the change in the tax refund policy and its related features need to be mapped into nodes and edges in the network, with nodes representing entities (such as goods, policies, tariff codes, etc.) and edges representing relationships between entities. On this basis, GNN will use the known relationships and data in the knowledge graph to perform graph propagation and information fusion, infer which areas may be affected, and identify new attributes or relationships. By constructing and training the model, the enterprise will obtain predicted information about the affected entities and adjust the entity relationships in the initial knowledge graph in real time. For example, if a change in a certain tax policy repeatedly affects a specific commodity, the enterprise can update the knowledge graph and its structure to ensure the timeliness and accuracy of the data. Finally, the knowledge graph will reflect the new relationships between all entities, becoming an important support for decision-making and operations, helping the enterprise to achieve effective compliance management in the constantly changing customs clearance environment. This updating process provides dynamic support for the enterprise to ensure synchronization with policies and efficient injection of information flow.
[0113] Reference Figure 3 The application also provides an intelligent tax refund declaration device for cross-border trade, which comprises:
[0114] The acquisition module 902 is configured to acquire a specified tax refund item and a plurality of tax refund data related to the specified tax refund item.
[0115] The extraction module 904 is configured to extract features of each of the tax refund data by using a preset adaptive parsing matrix, to obtain a plurality of commodity features.
[0116] The construction module 906 is configured to obtain and construct a mapping relationship database of commodity features and tariff codes based on each commodity feature by using a preset knowledge graph.
[0117] The filling module 908 is configured to extract declaration content in the mapping relationship database, and fill the declaration content into a preset declaration template to obtain a tax refund declaration form.
[0118] The tax refund module 910 is configured to perform a tax refund operation based on the tax refund declaration form.
[0119] In an embodiment, the intelligent tax refund declaration device for cross-border trade further comprises:
[0120] The data information recognition module is configured to recognize data information of each commodity to be traded by using a preset optical character recognition.
[0121] The data information analysis module is configured to recognize data information by using a preset natural language processing model to obtain each entity and the relationship between each entity.
[0122] The storage module is configured to store each entity as a node and store the relationship between each entity in a preset graph database to obtain the preset knowledge graph.
[0123] In an embodiment, the intelligent tax refund declaration device for cross-border trade further comprises:
[0124] The export country category extraction module is configured to perform semantic analysis on the specified tax refund item to extract an export country category.
[0125] The preset declaration template acquisition module is configured to acquire the preset declaration template from a preset template database based on the export country category.
[0126] In an embodiment, the intelligent tax refund declaration device for cross-border trade further comprises:
[0127] The multi-dimensional evaluation matrix construction module is configured to construct a multi-dimensional evaluation matrix for the tax refund declaration form based on the XGBoost algorithm.
[0128] The multi-dimensional evaluation matrix input module is configured to input the multi-dimensional evaluation matrix into a preset risk assessment model to obtain a corresponding risk level.
[0129] The review module is configured to send the tax refund declaration form to a specified terminal for review if the risk level is greater than a preset risk level.
[0130] In one embodiment, the intelligent tax refund declaration device for cross-border trade further comprises:
[0131] The customs inspection case acquisition module is configured to acquire a specified number of customs inspection cases from a preset case database and divide the specified number of customs inspection cases into a training data set and a verification data set according to a preset ratio;
[0132] The training module is configured to input the cases in the training data set into an initial risk assessment model for training to obtain a temporary risk assessment model;
[0133] The verification module is configured to verify the temporary risk assessment model by using the cases in the verification data set;
[0134] The marking module is configured to identify the temporary risk assessment model as the preset risk assessment model if the verification result is passed.
[0135] In one embodiment, the extraction module 904 comprises:
[0136] The data type acquisition submodule is configured to acquire the data type of each tax refund data;
[0137] The parsing matrix acquisition submodule is configured to acquire a corresponding parsing matrix based on the data type of each tax refund data;
[0138] The parsing matrix input submodule is configured to input the parsing matrix into a preset large language model to obtain a parsing model of the corresponding data type;
[0139] The tax refund data parsing submodule is configured to parse the corresponding tax refund data by using the parsing model of each data type to obtain a plurality of commodity features.
[0140] In one embodiment, the intelligent tax refund declaration device for cross-border trade further comprises:
[0141] The tax refund policy monitoring module is configured to monitor whether the tax refund policy changes in real time by using a preset crawler;
[0142] The text acquisition module is configured to acquire a first target text after the change and a second target text before the change if the change occurs;
[0143] The similarity calculation module is configured to calculate the similarity between the first target text and the second target text;
[0144] The similarity judgment module is configured to judge whether the similarity is greater than a threshold value;
[0145] The updating module is used to construct a GNN network to simulate the conduction effect if the similarity is greater than the threshold, obtain the predicted entities affected by the prediction, and update the relationship between the entities in the initial knowledge graph to obtain an updated preset knowledge graph.
[0146] Figure 4 FIG1 shows an internal structure diagram of a computer device in an embodiment. The computer device can be a terminal or a server. Figure 4 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement an intelligent tax refund declaration method for cross-border trade. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can implement an intelligent tax refund declaration method for cross-border trade. Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0147] In one embodiment, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:
[0148] Acquire a designated tax refund item and a plurality of tax refund data related to the designated tax refund item;
[0149] Extracting features from each of the tax refund data using a preset adaptive analytical matrix to obtain multiple product features;
[0150] Based on the characteristics of each product, a mapping relationship library between product characteristics and tariff codes is obtained and constructed through a preset knowledge graph;
[0151] Extracting the declaration content from the mapping relationship library and filling the declaration content into a preset declaration template to obtain a tax refund declaration form;
[0152] The tax refund operation is performed based on the tax refund declaration form.
[0153] The adaptive analysis matrix is used for feature extraction of the tax refund data, which can flexibly cope with different data types, optimize the structure and content of information, and provide high-quality input for subsequent analysis. Through the knowledge graph constructed based on commodity characteristics, the required declaration information can be quickly obtained, reducing the complexity and delay in the traditional tax refund declaration process, thereby improving the declaration efficiency, ensuring the consistency and accuracy of the data, effectively reducing the compliance risk caused by human errors, and enabling enterprises to perform tax refund operations more quickly and accurately, enhancing their competitiveness and flexibility in international trade.
[0154] In one embodiment, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the processor performs the following steps:
[0155] obtaining a specified tax refund item and a plurality of tax refund data related to the specified tax refund item;
[0156] extracting features of each of the tax refund data through a preset adaptive analysis matrix to obtain a plurality of commodity characteristics;
[0157] based on each commodity characteristic, obtaining and constructing a mapping relationship database of commodity characteristics and tariff codes through a preset knowledge graph;
[0158] extracting declaration content in the mapping relationship database and filling the declaration content into a preset declaration template to obtain a tax refund declaration form;
[0159] based on the tax refund declaration form, performing a tax refund operation.
[0160] The adaptive analysis matrix is used for feature extraction of the tax refund data, which can flexibly cope with different data types, optimize the structure and content of information, and provide high-quality input for subsequent analysis. Through the knowledge graph constructed based on commodity characteristics, the required declaration information can be quickly obtained, reducing the complexity and delay in the traditional tax refund declaration process, thereby improving the declaration efficiency, ensuring the consistency and accuracy of the data, effectively reducing the compliance risk caused by human errors, and enabling enterprises to perform tax refund operations more quickly and accurately, enhancing their competitiveness and flexibility in international trade.
[0161] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer readable storage medium, and when the program is executed, the processes of the above-mentioned embodiment methods can be included. Any reference to memory, storage, databases, or other media in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0162] The technical features of the above embodiments can be combined in any way. In order to make the description simple, not all possible combinations of the technical features in the above embodiments are described, but as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0163] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. An intelligent tax refund declaration method for cross-border trade, characterized by: The method comprises: Acquire a designated tax refund item and a plurality of tax refund data related to the designated tax refund item; Extracting features from each of the tax refund data using a preset adaptive analytical matrix to obtain multiple product features; Based on the characteristics of each product, a mapping relationship library between product characteristics and tariff codes is obtained and constructed through a preset knowledge graph; Extracting the declaration content from the mapping relationship library and filling the declaration content into a preset declaration template to obtain a tax refund declaration form; The tax refund operation is performed based on the tax refund declaration form.
2. The intelligent tax refund declaration method for cross-border trade according to claim 1 is characterized in that: Before the step of obtaining and constructing a mapping relationship library between product characteristics and tariff codes through a preset knowledge graph based on the characteristics of each product, the method further includes: Recognize the data information of each commodity to be traded through preset optical characters; Identify data information through a preset natural language processing model to obtain various entities and the relationships between them; Each entity is used as a node, and the relationship between each entity is stored in a preset graph database to obtain the preset knowledge graph.
3. The intelligent tax refund declaration method for cross-border trade according to claim 1 is characterized in that: Before the step of extracting the declaration content from the mapping relationship library and filling the declaration content into a preset declaration template to obtain a tax refund declaration form, the step further includes: Performing semantic analysis on the designated tax refund items to extract the exporting country category; Based on the exporting country category, the preset declaration template is obtained from a preset template database.
4. The intelligent tax refund declaration method for cross-border trade according to claim 1 is characterized in that: Before the step of performing the tax refund operation based on the tax refund declaration form, the method further includes: Constructing a multidimensional evaluation matrix for the tax refund declaration form based on the XGBoost algorithm; Inputting the multidimensional assessment matrix into a preset risk assessment model to obtain a corresponding risk level; If the risk level is greater than the preset risk level, the tax refund declaration form will be sent to a designated terminal for review.
5. The intelligent tax refund declaration method for cross-border trade according to claim 4 is characterized in that: Before the step of inputting the multidimensional assessment matrix into a preset risk assessment model to obtain a corresponding risk level, the method further includes: Obtain a specified number of customs audit cases from a preset case database, and divide the specified number of customs audit cases into a training data set and a validation data set according to a preset ratio; Inputting the cases in the training data set into the initial risk assessment model for training to obtain a temporary risk assessment model; Validating the temporary risk assessment model using cases in the validation dataset; If the verification result is passed, the temporary risk assessment model is identified as the preset risk assessment model.
6. The intelligent tax refund declaration method for cross-border trade according to claim 1 is characterized in that: The step of extracting features from each of the tax refund data using a preset adaptive analytical matrix to obtain multiple commodity features includes: Obtaining the data type of each tax refund data; Based on the data type of each tax refund data, obtaining a corresponding analytical matrix; Inputting the parsing matrix into a preset large language model to obtain a parsing model corresponding to the data type; The corresponding tax refund data is parsed using the parsing model of each data type to obtain multiple product features.
7. The intelligent tax refund declaration method for cross-border trade according to claim 1 is characterized in that: Before the step of obtaining and constructing a mapping relationship library between product characteristics and tariff codes through a preset knowledge graph based on the characteristics of each product, the method further includes: Monitor tax refund policies in real time using pre-set crawlers to see if there are any changes; If a change occurs, obtaining the first target text after the change and the second target text before the change; Calculating the similarity between the first target text and the second target text; Determining whether the similarity is greater than a threshold; If the similarity is greater than the threshold, a GNN network is constructed to simulate the conduction effect, the predicted entities affected are obtained, and the relationship between the entities is updated in the initial knowledge graph to obtain an updated preset knowledge graph.
8. An intelligent tax refund declaration device for cross-border trade, characterized in that: The device comprises: An acquisition module, configured to acquire a designated tax refund item and a plurality of tax refund data related to the designated tax refund item; An extraction module, configured to extract features from each of the tax refund data using a preset adaptive analytical matrix to obtain multiple commodity features; A construction module is used to obtain and construct a mapping relationship library between product characteristics and tariff codes based on the characteristics of each product through a preset knowledge graph; A filling module is used to extract the declaration content from the mapping relationship library and fill the declaration content into the preset declaration template to obtain a tax refund declaration form; The tax refund module is used to perform tax refund operations based on the tax refund declaration form.
9. A computer-readable storage medium, characterized in that A computer program is stored, and when the computer program is executed by a processor, the processor executes the steps of the intelligent tax refund declaration method for cross-border trade as claimed in any one of claims 1 to 7.
10. A computer device, characterized in that: The device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the intelligent tax refund declaration method for cross-border trade as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Cross-border tax system control method, device and system and readable storage medium
CN114154473A
Risk assessment method and device
CN115222405A
Internet big data analysis method and system based on artificial intelligence
CN119398824A
Method and device for automatically processing tax refund data, terminal and storage medium
CN119648447A
Data analysis method, device and equipment and computer readable storage medium
CN119849476A