Method and system for acquiring head-up information

Through multi-source data fusion and intelligent cleaning technology, combined with real-time offline hybrid acquisition mechanism, the problems of incomplete data, inaccuracy and insufficient timeliness in enterprise invoice systems are solved, efficient and reliable head-up information acquisition is achieved, and the coverage, accuracy and timeliness of the invoice system are improved.

CN120296072APending Publication Date: 2025-07-11WUXI BAISHANG ZHONGWANG DATA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510357818.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the enterprise invoice system has data dispersion, lag in updates, lack of multi-dimensional data fusion and intelligent cleaning capabilities, and cannot dynamically adapt to enterprise information changes, resulting in incomplete header information, poor accuracy and insufficient timeliness.

Method used

Multi-source data fusion, intelligent cleaning and real-time offline hybrid acquisition mechanisms are adopted, and the six elements of enterprise header information are extracted through natural language processing technology and deep learning models, format normalization and verification are carried out in combination with the rule database, and the optimal data is determined by similarity calculation and confidence scoring method, and data reception and update through real-time stream processing and offline batch processing.

Benefits of technology

It has achieved comprehensive coverage, high accuracy and real-time updates of enterprise header information, significantly improving the efficiency and reliability of the invoicing system, increasing the coverage rate by 300%, achieving accuracy of more than 99%, and improving the timeliness to minute levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296072A_ABST
    Figure CN120296072A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method and system for obtaining title information, and the method comprises the following steps: collecting enterprise title information, and carrying out the data receiving and updating through a real-time stream processing and offline batch processing mode, the enterprise title information comprising historical invoicing data, user input data and industrial and commercial data; extracting six elements of the enterprise head-up information by using a natural language processing technology and a deep learning model, and performing format normalization and verification on the six elements based on a rule base; combining multi-source data by adopting a similarity calculation method, and determining optimal data through a voting mechanism or a confidence scoring method; and performing classified storage on the data based on the confidence score, storing high-confidence data into a cache, storing low-confidence data into a recheck queue, and storing the data into a database after recheck. According to the embodiment of the invention, comprehensive coverage, high accuracy and real-time updating of the title information can be realized, and the efficiency and reliability of the billing system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and particularly relates to a method and system for obtaining header information. Background Art

[0002] In the enterprise invoicing process, the integrity, accuracy, and timeliness of header information are crucial for tax compliance and financial management. Existing technologies mostly rely on a single data source (such as self-reported by enterprises or third-party static databases), and have the following limitations:

[0003] 1. Data is scattered and updated lagged: There is no linkage between different data sources, and the update period is long (usually ≥ 24 hours), unable to reflect enterprise information changes in real time.

[0004] 2. Lack of multi-dimensional data fusion and intelligent cleaning capabilities: Existing solutions cannot effectively integrate multi-source data, and lack intelligent means to clean and verify data, resulting in a high error rate.

[0005] 3. Unable to dynamically adapt to enterprise information changes: Traditional systems rely on offline batch updates and cannot track dynamic changes in enterprise business information, bank accounts, etc. in real time, resulting in expired invoice information.

[0006] Application Content

[0007] The purpose of the embodiments of this application is to provide a method and system for obtaining header information to solve the defect of high error rate in the prior art.

[0008] To solve the above technical problems, this application is implemented as follows:

[0009] In the first aspect, a method for obtaining header information is provided, including the following steps:

[0010] Collect enterprise header information, and perform data reception and update through real-time stream processing and offline batch processing. The enterprise header information includes historical invoicing data, user input data, and business data;

[0011] Use natural language processing technology and deep learning models to extract six elements of the enterprise header information, and perform format normalization and verification on the six elements based on a rule library;

[0012] Adopt a similarity calculation method to merge multi-source data, and determine the optimal data through a voting mechanism or a confidence score method;

[0013] Classify and store data based on confidence scores, store high-confidence data in the cache, store low-confidence data in the review queue, and store it in the database after review.

[0014] In the second aspect, a system for obtaining header information is provided, including:

[0015] A collection module, configured to collect enterprise header information, and perform data reception and update through real-time stream processing and offline batch processing, where the enterprise header information includes historical invoicing data, user input data, and industrial and commercial data;

[0016] An extraction module, configured to extract six elements of the enterprise header information by using natural language processing technology and a deep learning model, and perform format normalization and verification on the six elements based on a rule library;

[0017] A merging module, configured to merge multi-source data by using a similarity calculation method, and determine optimal data through a voting mechanism or a confidence scoring method;

[0018] A storage module, configured to classify and store data based on a confidence score, store high-confidence data in a cache, store low-confidence data in a review queue, and store it in a database after review.

[0019] The embodiments of the present application adopt multi-source data fusion, intelligent cleaning and verification, and real-time and offline hybrid collection to achieve comprehensive coverage, high accuracy, and real-time update of header information, and can improve the efficiency and reliability of the invoicing system. Description of the Drawings

[0020] Figure 1 is a flowchart of a method for obtaining header information provided by an embodiment of the present application;

[0021] Figure 2 is a system architecture diagram of a method for obtaining header information provided by an embodiment of the present application;

[0022] Figure 3 is a specific implementation diagram of a method for obtaining header information provided by an embodiment of the present application;

[0023] Figure 4 is a structural schematic diagram of a system for obtaining header information provided by an embodiment of the present application. Detailed Embodiments

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0025] Currently, the invoicing system combines industrial and commercial data APIs with enterprise historical invoicing records to achieve automatic filling of tax numbers, but there are the following limitations:

[0026] 1. Limited coverage: It only covers the enterprise name and tax number, lacking four elements such as phone number, address, bank of deposit, etc., and the information is incomplete.

[0027] 2. Long update cycle: The update of industrial and commercial data depends on batch synchronization, with a cycle ≥ 24 hours, unable to meet the real-time requirements.

[0028] 3. Data conflicts unresolved: When data from different sources is inconsistent (such as different industrial and commercial data from the user-entered address), the system cannot intelligently resolve conflicts and relies on manual intervention.

[0029] In view of the above problems, the embodiments of the present application aim to comprehensively improve the performance of the invoicing system through the following technical means:

[0030] 1. Multi-source data fusion: Integrate the three core data sources of the data warehouse invoice pool, invoicing client, and industrial and commercial data, covering six elements of the enterprise header information (name, phone number, address, tax number, bank of deposit address, account number) to ensure the comprehensiveness of information.

[0031] 2. Intelligent verification mechanism: Combine deep learning models (such as BERT), natural language processing (NLP) technology, and rule engines to achieve intelligent data cleaning, semantic matching, and logical verification, and improve the data accuracy to over 99%.

[0032] 3. Dynamic update ability: Adopt a real-time and offline hybrid acquisition mechanism (Kafka real-time stream + Spark offline batch processing) and an incremental update algorithm (based on timestamp differences) to achieve minute-level information synchronization and ensure data timeliness. Among them, the incremental update hash algorithm uses SHA-256 to compare different fields.

[0033] Through the above innovations, the embodiments of the present application solve the pain points of the traditional invoicing system in terms of coverage, accuracy, and timeliness. Through multi-source data fusion, intelligent cleaning, and real-time update mechanisms, the above problems are comprehensively solved, and the coverage, accuracy, and timeliness of the invoicing system are improved, providing efficient and reliable invoicing services for enterprises.

[0034] Next, in combination with the accompanying drawings, a method for obtaining header information provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0035] As Figure 1 shown, it is a flowchart of a method for obtaining header information provided by the embodiments of the present application. The method includes the following steps:

[0036] Step 101, collect enterprise header information, and receive and update data through real-time stream processing and offline batch processing. The enterprise header information includes historical invoicing data, user input data, and industrial and commercial data.

[0037] Step 102: Extract the six elements of the enterprise header information using natural language processing technology and deep learning models, and perform format normalization and verification on the six elements based on a rule library.

[0038] Specifically, a preset address standardization rule library can be adopted and combined with a BERT model for semantic recognition and field matching.

[0039] Among them, the address standardization rule library contains a standardized mapping table for administrative divisions, road names, and building numbers, and automatically converts non-standard addresses based on a string matching algorithm.

[0040] Step 103: Merge multi-source data using a similarity calculation method, and determine the optimal data through a voting mechanism or a confidence scoring method.

[0041] Specifically, a cosine similarity algorithm can be used to calculate the matching degrees of enterprise names, addresses, and tax numbers, and a weighted voting mechanism can be combined to select the optimal data source. The weights of different data sources are set according to the reliability of the data sources.

[0042] Among them, the weight of industrial and commercial data is greater than the weight of the historical invoicing data, and the weight of the historical invoicing data is greater than the weight of the user input data.

[0043] Step 104: Classify and store the data based on the confidence score, store the high-confidence data in the cache, and store the low-confidence data in the review queue and then in the database after review.

[0044] The embodiments of this application adopt multi-source data fusion, intelligent cleaning and verification, and real-time and offline hybrid collection to achieve comprehensive coverage, high accuracy, and real-time update of header information, and can improve the efficiency and reliability of the invoicing system.

[0045] In the embodiments of this application, as Figure 2 shown, the system architecture is divided into four layers (data collection layer, data processing layer, verification and storage layer, and service layer). Through multi-source data fusion, intelligent processing, and efficient storage and invocation, the problems of coverage, accuracy, and timeliness of traditional invoicing systems are comprehensively solved:

[0046] Among them, the data collection layer is used to integrate three core data sources:

[0047] 1. Data warehouse invoice pool: Load offline historical invoicing data and provide enterprise historical header information.

[0048] 2. Invoicing client: Receive the enterprise header information input by the user in real time to ensure data real-time.

[0049] 3. Business data: Incrementally crawl enterprise registration information through a distributed crawler framework (such as Scrapy-Redis) and dynamically update the data.

[0050] The data collection layer adopts a real-time and offline hybrid collection mechanism, taking into account the depth of historical data and the requirements for real-time data updates.

[0051] The data processing layer is used to solve multi-source data conflicts and includes the following modules:

[0052] 1. Multi-source data fusion module: Deduplicate and merge multi-source data based on a similarity algorithm (such as cosine similarity) to solve the data conflict problem.

[0053] 2. NLP cleaning module: Use the BERT model to extract key fields (such as address, enterprise name) and perform standardization processing (such as address normalization).

[0054] 3. Intelligent scoring module: Generate a data confidence score through Embedding vectorization and weight matrix calculation to screen high-quality data.

[0055] The data processing layer combines AI and a rule engine to ensure high data accuracy (accuracy ≥ 99%).

[0056] The verification and storage layer is used for efficient storage and fast response:

[0057] 1. Rule engine verification: Perform logical verification on the six elements (such as tax number verification digit, address administrative division matching) to ensure data compliance.

[0058] 2. Distributed storage: Persist high-score data to a MongoDB sharded cluster to support massive data storage and efficient query; store low-score data in a Redis queue for re-verification, triggering the manual intervention process.

[0059] The service layer is used to provide low-latency API interfaces:

[0060] 1. Provide low-latency API interfaces for the invoicing system to call enterprise header information in real time, with a response time ≤ 50ms, meeting the requirements of high-concurrency business.

[0061] Through the above four-layer architecture, the embodiments of the present application achieve comprehensive coverage, high accuracy, and real-time update of enterprise header information, significantly improving the efficiency and reliability of the invoicing system, and solving the pain points of traditional invoicing systems in terms of coverage, accuracy, and timeliness.

[0062] Furthermore, as Figure 3 shown, the process implementation method of the present application is divided into four core links:

[0063] Data collection: Receive the input from the invoicing client through Kafka real-time stream to ensure data timeliness; Use Spark offline batch processing to load historical data in the data warehouse to provide historical information support; Adopt an incremental crawling algorithm (based on timestamp difference) to dynamically crawl industrial and commercial data to achieve minute-level updates.

[0064] Specifically, it includes the following steps:

[0065] 1. Real-time stream processing: The input from the invoicing client is transmitted through Kafka real-time stream to ensure data timeliness, and the response time ≤ 50ms.

[0066] 2. Offline batch processing: Use Spark SQL to load historical invoice data (in Parquet / ORC format) in the data warehouse, and extract six-element fields by grouping according to enterprise ID.

[0067] 3. Incremental crawling: Incrementally crawl industrial and commercial data through a distributed crawler framework (such as Scrapy-Redis), and reduce 90% of duplicate data crawling based on the timestamp difference algorithm.

[0068] Data cleaning: Use an NLP model (such as BERT) to identify address elements, and implement format normalization based on a rule library (such as 'Beijing Ding District' → 'Haidian District' → 'Jinghai'); Solve multi-source data conflicts through a voting mechanism, and select the field with the highest frequency as the final value.

[0069] Specifically, it includes the following steps:

[0070] 1. NLP normalization: The BERT model identifies entities such as administrative regions, streets, and house numbers in the address, and standardizes the address through a built-in rule library (such as 'Beijing' → 'Jing').

[0071] 2. Conflict resolution: Count the occurrence frequencies of each field in multi-source data, and select the value with the highest frequency as the final result (adopt it if it appears ≥ 2 times in industrial and commercial data).

[0072] Intelligent scoring: Use Embedding vectors to calculate field similarity and identify semantic consistency; Generate a comprehensive score based on a weight matrix (industrial and commercial data 0.6, data warehouse data 0.3, client 0.1) to quantify data credibility.

[0073] Specifically, it includes the following steps:

[0074] 1. Embedding vectorization: Use the Sentence-BERT model to convert texts such as enterprise names and addresses into 768-dimensional vectors, and calculate the cosine similarity (if the similarity between 'Tencent Technology' and 'Tencent Technology Co., Ltd.' ≥ 0.9, it is determined to be the same enterprise).

[0075] 2. Weight Matrix Scoring: The comprehensive scoring formula is Comprehensive Score = Σ(Field Confidence × Data Source Weight). Example: If the enterprise name is matched in the industrial and commercial data (confidence 1.0), then the score is 1.0 × 0.6 = 0.6.

[0076] Storage and Call: High-score data (score ≥ 90) is stored in the Redis cache to support low-latency queries; low-score data triggers the manual review process and is updated to the database after ensuring data accuracy.

[0077] Specifically, it includes the following steps:

[0078] 1. High-score Data Storage: Data with a score ≥ 90 is stored in the Redis cache (Hash table structure, Key is the enterprise name, Value is the six elements in JSON format), and the response time ≤ 50ms.

[0079] 2. Low-score Data Review: Data with a score < 90 is written into the Redis review queue (List structure), and the background management system pushes it to the manual review interface. After confirmation, it is updated to MongoDB and synchronized to the Redis cache.

[0080] The embodiments of this application have the following innovation points:

[0081] Multi-source Data Fusion Method: Real-time and offline hybrid acquisition mechanism, integrating data warehouse, client, and industrial and commercial data, covering the six elements of the enterprise; incremental crawling algorithm (based on timestamp difference) to dynamically update industrial and commercial data.

[0082] Intelligent Data Processing Technology: Use the BERT model to extract key fields and perform standardization processing; generate data confidence scores through Embedding vectorization and weight matrix calculation.

[0083] Dynamic Update and Efficient Storage Solution:

[0084] 1. High-score Data Storage Method: Store data with a score ≥ 90 in the Redis cache to support low-latency queries.

[0085] 2. Low-score Data Review Method: Write data with a score < 90 into the Redis review queue to trigger the manual review process.

[0086] Rule Engine and AI Collaborative Verification Technology:

[0087] 1. Rule Engine Verification Method: Perform logical verification on the six elements (such as tax number verification digit, address administrative division matching).

[0088] 2. AI and Rule Engine Collaborative Verification Method: Combine the AI model and the rule engine to ensure high data accuracy.

[0089] Compared with the prior art, the embodiments of the present application have the following advantages:

[0090] 1. Higher coverage

[0091] Integrate the three core data sources of the data warehouse invoice pool, the invoicing client, and industrial and commercial data, comprehensively covering the six enterprise elements (name, phone number, address, tax number, bank address, account number). Through multi-source data fusion technology, solve the problem of data dispersion, ensure information integrity, increase the coverage rate by 300%, and meet the requirements of tax compliance and user experience.

[0092] 2. Higher accuracy

[0093] Introduce deep learning models (such as BERT) and natural language processing (NLP) technologies to intelligently clean and standardize data. Combine rule engines (such as tax number check digits, address administrative division matching) with AI models to ensure high data accuracy (accuracy rate ≥ 99%). The accuracy is significantly improved, reducing invoice returns or cancellations and lowering the operating costs of enterprises.

[0094] 3. Stronger timeliness

[0095] Adopt a real-time and offline hybrid acquisition mechanism (Kafka real-time stream + Spark offline batch processing) to balance the needs of in-depth historical data and real-time data updates. The incremental crawling algorithm (based on timestamp differences) dynamically crawls industrial and commercial data to achieve minute-level updates. The timeliness of data updates is improved to the minute level, ensuring the real-time and accurate enterprise information. After a certain enterprise is connected, the invoicing error rate drops from 7% to 0.3%, and the processing efficiency is increased by 20 times.

[0096] 4. Higher degree of intelligence

[0097] Through Embedding vectorization and weight matrix calculation, generate data confidence scores to quantify the credibility of data. High-score data is automatically stored, and low-score data triggers manual review to balance automation and manual intervention. The degree of intelligence is significantly improved, reducing manual intervention and improving system efficiency.

[0098] 5. Higher storage and call efficiency

[0099] High-score data is stored in MongoDB and synchronized to the Redis cache to support low-latency queries (response time ≤ 50ms). Low-score data triggers the manual review process, and after ensuring data accuracy, it is updated to the database. The storage and call efficiency is significantly improved, supporting high-concurrency business requirements and ensuring system stability.

[0100] As Figure 4 shown, it is a schematic structural diagram of a system for obtaining header information provided by the embodiments of the present application, including:

[0101] The acquisition module 410 is used to acquire enterprise name information, and receive and update data through real-time stream processing and offline batch processing. The enterprise name information includes historical invoicing data, user input data, and industrial and commercial data.

[0102] The extraction module 420 is used to extract six elements of the enterprise name information by using natural language processing technology and deep learning models, and normalize and verify the formats of the six elements based on a rule library.

[0103] Specifically, the extraction module 420 is specifically used to adopt a preset address standardization rule library and combine it with a BERT model for semantic recognition and field matching.

[0104] Among them, the address standardization rule library contains a standardization mapping table for administrative regions, road names, and building numbers, and automatically converts non-standard addresses based on a string matching algorithm.

[0105] The merging module 430 is used to merge multi-source data by using a similarity calculation method, and determine the optimal data through a voting mechanism or a confidence score method.

[0106] Specifically, the merging module 430 is specifically used to calculate the matching degrees of enterprise names, addresses, and tax numbers by using a cosine similarity algorithm, and combine a weighted voting mechanism to select the optimal data source. The weights of different data sources are set according to the reliability of the data sources.

[0107] Among them, the weight of industrial and commercial data is greater than the weight of the historical invoicing data, and the weight of the historical invoicing data is greater than the weight of the user input data.

[0108] The storage module 440 is used to classify and store data based on the confidence score, store high-confidence data in the cache, store low-confidence data in the review queue, and store them in the database after review.

[0109] The embodiments of the present application adopt multi-source data fusion, intelligent cleaning and verification, and real-time and offline hybrid acquisition to achieve comprehensive coverage, high accuracy, and real-time update of the name information, and can improve the efficiency and reliability of the invoicing system.

[0110] The embodiments of the present application also provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it realizes each process of the method embodiment for obtaining the name information described above and can achieve the same technical effects. To avoid repetition, it will not be described in detail here. Among them, the computer-readable storage medium is, for example, a read-only memory (Read-Only Memory, abbreviated as ROM), a random access memory (Random Access Memory, abbreviated as RAM), a magnetic disk, or an optical disc, etc.

[0111] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0112] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on this understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0113] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A method for obtaining header information, characterized in that It includes the following steps: Collect the enterprise header information, and receive and update the data through real-time stream processing and offline batch processing. The enterprise header information includes historical invoicing data, user input data, and industrial and commercial data; Use natural language processing technology and deep learning models to extract the six elements of the enterprise header information, and perform format normalization and verification on the six elements based on a rule library; Adopt a similarity calculation method to merge multi-source data, and determine the optimal data through a voting mechanism or a confidence scoring method; Classify and store the data based on the confidence score, store the high-confidence data in the cache, store the low-confidence data in the review queue, and store it in the database after review.

2. The method according to claim 1, characterized in that, The format normalization and verification of the six elements based on the rule library specifically include: Adopt a preset address standardization rule library and combine it with the BERT model for semantic recognition and field matching.

3. The method according to claim 2, wherein The address standardization rule library contains a standardization mapping table of administrative regions, road names, and building numbers, and automatically converts non-standard addresses based on a string matching algorithm.

4. The method according to claim 1, wherein The adoption of a similarity calculation method to merge multi-source data and determine the optimal data through a voting mechanism or a confidence scoring method specifically includes: Adopt the cosine similarity algorithm to calculate the matching degrees of enterprise names, addresses, and tax numbers, and combine a weighted voting mechanism to select the optimal data source. The weights of different data sources are set according to the reliability of the data sources.

5. The method according to claim 4, characterized in that, The weight of the industrial and commercial data is greater than the weight of the historical invoicing data, and the weight of the historical invoicing data is greater than the weight of the user input data.

6. A system for obtaining header information, characterized in that, It includes: A collection module for collecting the enterprise header information, and receiving and updating the data through real-time stream processing and offline batch processing. The enterprise header information includes historical invoicing data, user input data, and industrial and commercial data; An extraction module for using natural language processing technology and deep learning models to extract the six elements of the enterprise header information, and performing format normalization and verification on the six elements based on a rule library; A merging module for adopting a similarity calculation method to merge multi-source data, and determining the optimal data through a voting mechanism or a confidence scoring method; A storage module for classifying and storing the data based on the confidence score, storing the high-confidence data in the cache, storing the low-confidence data in the review queue, and storing it in the database after review.

7. The system according to claim 6, wherein The extraction module is specifically used for adopting a preset address standardization rule library and combining it with the BERT model for semantic recognition and field matching.

8. The system according to claim 7, wherein The address standardization rule library contains a standardization mapping table of administrative regions, road names, and building numbers, and automatically converts non-standard addresses based on a string matching algorithm.

9. The system according to claim 6, wherein The merging module is specifically used for adopting the cosine similarity algorithm to calculate the matching degrees of enterprise names, addresses, and tax numbers, and combining a weighted voting mechanism to select the optimal data source. The weights of different data sources are set according to the reliability of the data sources.

10. The system according to claim 9, wherein The weight of the industrial and commercial data is greater than the weight of the historical invoicing data, and the weight of the historical invoicing data is greater than the weight of the user input data.