Document table processing system based on question and answer reasoning

Through a document form processing system based on question-answer reasoning, natural language processing and machine learning algorithms are used to solve the problems of inefficiency in traditional data processing and prone to human errors, realizing in-depth semantic understanding and data analysis, and improving the ease of use and operational efficiency of the system.

CN119962683AInactive Publication Date: 2025-05-09SHANGHAI BINGGE NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510041422.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional document form data processing methods require users to have professional skills, are inefficient and prone to human errors. Especially when the data volume and complexity are large, intelligent processing systems are lacking to quickly and accurately analyze user needs.

Method used

Provide a document form processing system based on question-answer reasoning, including data acquisition and integration layer, question-answer understanding and transformation layer, reasoning and analysis layer, and result presentation and interaction layer, using natural language processing, rules engines and machine learning algorithms to achieve deep semantic understanding and data analysis.

Benefits of technology

Realize deep semantic understanding through natural language processing technology to improve the accessibility and ease of use of the system; use machine learning algorithms for data analysis, provide accurate trend prediction, customer segmentation and market positioning, optimize resource allocation and improve operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962683A_ABST
    Figure CN119962683A_ABST
Patent Text Reader

Abstract

The invention provides a document table processing system based on question and answer reasoning, and relates to the technical field of data processing and analysis, and the system comprises a data collection and integration layer which is used for collecting document table data of an enterprise and carrying out the preprocessing operation; the question and answer understanding and conversion layer is integrated with a natural language processing engine to perform semantic analysis processing on questions proposed by the user; the reasoning and analyzing layer is used for designing a reasoning mechanism based on a rule engine and a machine learning algorithm; a machine learning algorithm is used for pattern recognition, trend prediction and anomaly detection; and the result presentation and interaction layer adopts a data visualization technology and a report generation tool according to a result type and a user demand. According to the method, deep semantic understanding is realized by using a natural language processing technology of the word vector model and the syntactic analyzer, so that a user can express complex data requirements by using a daily language, and the accessibility and usability of the system are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing and analysis, and in particular to a document form processing system based on question-answering reasoning. Background Art

[0002] In many key areas such as corporate management, financial analysis, and academic research, document and form data undoubtedly carry extremely important and rich information treasures. Taking corporate management as an example, financial statements are a direct reflection of the company's operating conditions. The balance sheet lists in detail the distribution of the company's assets, liabilities, and owner's equity. The income statement accurately presents the company's profit level and income and expenditure structure during the period, and the cash flow statement tracks and records the inflow and outflow paths of the company's funds.

[0003] However, traditional processing methods often require users to have professional data processing skills to manually extract and analyze data from tables, which is not only inefficient but also prone to human errors. With the continuous increase in data volume and data complexity, there is an urgent need for a system that can intelligently process document table data, accurately understand user needs and quickly provide analysis results. Therefore, it is necessary to provide a document table processing system based on question-answering reasoning to solve the above technical problems. Summary of the invention

[0004] The present invention provides a document form processing system based on question-answering reasoning, which solves the problems raised in the above-mentioned background technology.

[0005] In order to solve the above technical problems, the document form processing system based on question-answering reasoning provided by the present invention comprises:

[0006] The data collection and integration layer is used to collect the company's document form data and perform pre-processing operations and data storage;

[0007] The question-answer understanding and conversion layer integrates a natural language processing engine to perform semantic analysis on the questions raised by users and sends the analysis results after semantic analysis to the reasoning and analysis layer; the natural language processing engine includes a word vector model and a syntactic analyzer based on deep learning;

[0008] Reasoning and analysis layer: design reasoning mechanism based on rule engine and machine learning algorithm. The rule engine predefines business rules and logical relationships. Use machine learning algorithms for pattern recognition, trend prediction and anomaly detection. Machine learning algorithms include regression analysis and cluster analysis. The results obtained by machine learning algorithms are passed to the result presentation and interaction layer.

[0009] The result presentation and interaction layer uses data visualization technology and report generation tools to present the results in the form of intuitive charts, text reports or interactive data dashboards according to the result type and user needs.

[0010] As a preferred method, based on a natural language processing engine, a word vector model is used to perform semantic analysis on the questions raised by the user, specifically:

[0011] Obtain a trained word vector model, use the word vector model to identify the words in the questions raised by the user and convert them into word vectors; adopt a neural network architecture based on the self-attention mechanism, and learn the contextual relationship of the words in the user's questions through a multi-layer self-attention mechanism;

[0012] Among them, the calculation formula of the self-attention mechanism in the neural network architecture is as follows:

[0013] For the input sequence X=(x1,x2,…x H ), where H represents the total number of elements in the input sequence X; the query vector Q, key vector K and value vector V are obtained through linear transformation; the formula is: Q = XW Q , K = XW K , V = XW V ;W Q , W K , W V Represents the learnable weight matrix; then calculate its attention score S, the formula is expressed Among them, d k represents the dimension of the key vector K, QK T Matrix multiplication representing the transpose of the query vector and the key vector;

[0014] Use the Softmax function to normalize the attention score and value vector to obtain the attention weight O, and the formula is O=softmax(S)V.

[0015] Preferably, based on a natural language processing engine, a syntax analyzer is used to perform syntax analysis on the user's question, and the specific steps are as follows:

[0016] Determine the sentence grammatical structure and identify the non-terminal symbol N in the user's question r With the terminal symbol T v , according to the probabilistic context-free grammar, there is a production rule N r →β, where β represents a string of symbols consisting of non-terminal symbols and terminal symbols, and v represents the index of different terminal symbols in the sentence;

[0017] For the analysis tree T of any sentence in the user's question, its probability P(T) is obtained by multiplying the probabilities of all production rules, and the formula is expressed as: Among them, Nr→βr represents the production rules in the analysis tree T; r represents the index variable of different symbols in the sentence, marking the vocabulary vector and grammatical structure in the user's question as the question analysis result;

[0018] Build the semantic knowledge base of the enterprise, including enterprise management, financial analysis, academic research and their corresponding professional terms, concept relationships and business rules; match and map the problem analysis results of the user's questions with the entries in the semantic knowledge base, calculate the semantic similarity between the problem analysis results and any knowledge base entry in the semantic knowledge base, obtain the keyword vector in the user's question represented by A, identify the entry vector in the semantic knowledge base represented by B, use the cosine similarity algorithm to measure the similarity between the keywords in the user's question and the semantic knowledge base entries, marked as semantic similarity, the cosine similarity algorithm formula is expressed as: The knowledge base entry with the largest semantic similarity is recorded as a semantically related entry; where Ai represents the i-th component of the user question keyword vector A, Bi represents the i-th component of the semantic knowledge base entry vector B, and n represents the dimension of the keyword vector A and the semantic knowledge base entry vector B, that is, the number of elements in the vector; all semantic similarities and their corresponding target knowledge base entries and semantically related entries are recorded as semantic matching results;

[0019] Determine the semantically related items, and parse the problem requirements based on the connotation of the semantically related items and the business rules and data processing logic associated with them; obtain the pre-set data storage structure and processing method, and build the corresponding structured query language or data processing task instruction template; fill the specific parameters of the semantically related items in the questions raised by the user into the structured query language or data processing task instruction template to generate a structured query language or data processing task instruction; send the semantic matching results and the structured query language or data processing task instruction to the reasoning and analysis layer.

[0020] As a preferred method, based on the reasoning and analysis layer, the questions raised by the user are analyzed by regression analysis in the machine learning algorithm, and the specific analysis steps are as follows:

[0021] According to the rule engine, the questions raised by the user are identified and analyzed to determine the independent variable x and the dependent variable y; the linear regression model is selected, and the formula is y=B0+B1x+ε, where B0 and B1 represent regression coefficients and ε represents the error term;

[0022] The regression coefficients are calculated by the least squares method, and B0 and B1 are solved by minimizing the sum of squared errors; the formula is: c represents the number of data points, xt and yt represent the values ​​of the independent variable and dependent variable of the tth data point, respectively. are the means of the independent and dependent variables, respectively.

[0023] As a preferred method, based on the reasoning and analysis layer, cluster analysis in a machine learning algorithm is used to analyze the questions raised by the user, and the specific analysis steps are as follows:

[0024] Set the number of clusters P;

[0025] Initialize P cluster centers and randomly select P data points as initial cluster centers;

[0026] The Euclidean distance formula is used to calculate the distance from each data point to the cluster center. The formula is: Among them, xtp represents the value of the t-th data point in the p-th dimension, swp represents the value of the w-th cluster center in the p-th dimension, and m represents the data dimension;

[0027] Assign the data point to the class with the closest cluster center;

[0028] Then recalculate the cluster center and convergence judgment, calculate the average value of the data points in each class as the new cluster center; calculate the difference between the previous cluster center and the current cluster center to get the cluster difference;

[0029] The number of iterations of the current cluster center is counted, and the cluster difference values ​​of all cluster centers and the number of iterations are weighted to obtain the cluster difference value; a cluster difference threshold is set. If the cluster difference value is less than the cluster difference threshold, it means that the cluster center no longer changes or the change range is within a controllable range, and the clustering process converges.

[0030] Preferably, based on the number of clusters P, the elbow rule is used to determine the optimal number of clusters, and the specific steps are:

[0031] The sum of squares of each data point and its cluster center is calculated and marked as the sum value SSE. The formula is: Where Q represents the number of cluster centers, Cw represents the w-th cluster center, x represents the data point, and uw represents the center of the w-th cluster;

[0032] Starting from P=2, increase P by one in each round; for each P value, calculate the corresponding aggregation value of each P value; construct a relationship line graph with P value as the horizontal axis and aggregation value as the vertical axis, mark the position of the aggregation value in the relationship line graph as the aggregation point, connect adjacent aggregation points, calculate the slope between adjacent aggregation points, and the formula is g(P)=SSE(P+1)-SSE(P); then calculate the rate of change of its slope Δg(P)=g(P+1)-g(P);

[0033] A change threshold is set. If the change rate is less than the change threshold, the P value corresponding to the change rate is used as the optimal number of clusters.

[0034] As a preferred embodiment, based on the data collection and integration layer, the frequency of data collection is dynamically adjusted according to the real-time requirements of enterprise data update, specifically:

[0035] Establish a data update monitoring module to monitor the update timestamp or data change mark of the data in the data source in real time;

[0036] Set a data update time zone, identify the number of times data is updated in the data update time zone and mark it as a number of times; divide the number of times by the duration corresponding to the data update time zone to obtain the data update frequency f;

[0037] Set the normal update frequency range, match the data update frequency with the normal update frequency range, if the data update frequency is within its normal update frequency range, keep the current data collection frequency unchanged;

[0038] Identify the upper limit value T1 and the lower limit value T2 of the normal update frequency range,

[0039] If the data update frequency is greater than or equal to the upper limit of its normal update frequency range, it means that the data is updated frequently. Then the dynamic data acquisition adjustment coefficient is calculated and the formula is expressed as: Use the dynamic data acquisition adjustment coefficient b1 to adjust the data acquisition frequency; if the data update frequency is less than or equal to the lower limit of its normal update frequency range, it means that the data update is not frequent, then calculate and reduce the dynamic data acquisition adjustment coefficient Use the dynamic data acquisition adjustment factor b2 to adjust the data acquisition frequency.

[0040] Preferably, based on the data collection and integration layer, the preprocessing operation includes cleaning, conversion, and format normalization operations; the document form data after the preprocessing operation is stored in a distributed data repository in a unified data storage format and model.

[0041] Compared with the related art, the document form processing system based on question-answering reasoning provided by the present invention has the following beneficial effects:

[0042] 1. The present invention achieves deep semantic understanding by utilizing the natural language processing technology of word vector models and syntactic analyzers, allowing users to express complex data requirements in everyday language, greatly improving the accessibility and ease of use of the system.

[0043] 2. The present invention uses machine learning algorithms to deeply mine the complex relationships and potential patterns inherent in data, uses regression analysis to accurately predict trends, and provides forward-looking guidance for corporate strategic planning; cluster analysis effectively identifies data distribution characteristics and achieves accurate customer segmentation and market positioning; forms a system that uses regression analysis to assist strategic foresight and cluster analysis to achieve precise marketing, thereby optimizing resource allocation, driving innovative development, and improving operational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A schematic diagram of the principles of the document form processing system based on question-answering reasoning provided by the present invention. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] The terms used in this disclosure are for the purpose of describing embodiments only and are not intended to limit the disclosure. The singular forms of "group", "class" and "the" used in this disclosure and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0047] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0048] Please refer to Figure 1 The document form processing system based on question-answering reasoning includes:

[0049] The data collection and integration layer is used to collect the company's document form data and perform pre-processing operations and data storage;

[0050] The question-answer understanding and conversion layer integrates a natural language processing engine to perform semantic analysis on the questions raised by users and sends the analysis results after semantic analysis to the reasoning and analysis layer; the natural language processing engine includes a word vector model and a syntactic analyzer based on deep learning;

[0051] The reasoning and analysis layer designs a reasoning mechanism based on the rule engine and machine learning algorithms. The rule engine predefines business rules and logical relationships. It should be noted that the business rules are specifically formulated based on professional knowledge in the fields of enterprise management, financial analysis, academic research, etc.; machine learning algorithms are used for pattern recognition, trend prediction, and anomaly detection; machine learning algorithms include regression analysis and cluster analysis; the results obtained by the machine learning algorithm are passed to the result presentation and interaction layer;

[0052] The result presentation and interaction layer uses data visualization technology and report generation tools to present the results in the form of intuitive charts, text reports or interactive data dashboards according to the result type and user needs.

[0053] In this application, based on the natural language processing engine, the word vector model is used to perform semantic analysis on the questions raised by the user, specifically:

[0054] Obtain a trained word vector model, use the word vector model to identify the words in the questions raised by the user and convert them into word vectors; adopt a neural network architecture based on a self-attention mechanism, and learn the contextual relationship of the words in the user's questions through a multi-layer self-attention mechanism; it should be noted that in actual applications, the training data comes from a large-scale text corpus, such as news articles, academic documents, corporate documents, etc.; it is considered a mature technology, so its training method is not clearly proposed in this application;

[0055] Among them, the calculation formula of the self-attention mechanism in the neural network architecture is as follows:

[0056] For the input sequence X=(x1,x2,…x H ), where H represents the total number of elements in the input sequence X; the query vector Q, key vector K and value vector V are obtained through linear transformation; the formula is: Q = XW Q , K = XW K , V = XW V ;W Q , W K , W V Represents the learnable weight matrix; then calculate its attention score S, the formula is expressed Among them, d k represents the dimension of the key vector K, QK T Matrix multiplication representing the transpose of the query vector and the key vector;

[0057] Use the Softmax function to normalize the attention score and value vector to obtain the attention weight O, and the formula is O=softmax(S)V.

[0058] In this application, based on the natural language processing engine, a syntactic analyzer is used to perform syntactic analysis on user questions, and the specific steps are as follows:

[0059] Determine the sentence grammatical structure and identify the non-terminal symbol N in the user's question i With the terminal symbol T j , according to the probabilistic context-free grammar, there is a production rule N r →β, where β represents a symbol string consisting of non-terminal symbols and terminal symbols;

[0060] For the analysis tree T of any sentence in the user's question, its probability P(T) is obtained by multiplying the probabilities of all production rules, and the formula is expressed as: Among them, Nr→βr represents the production rules in the analysis tree T; r represents the index variable of different symbols in the sentence;

[0061] Mark the vocabulary vectors and grammatical structures in the user's questions as question analysis results;

[0062] Build the semantic knowledge base of the enterprise, including enterprise management, financial analysis, academic research and their corresponding professional terms, concept relationships and business rules; match and map the problem analysis results of the user's questions with the entries in the semantic knowledge base, calculate the semantic similarity between the problem analysis results and any knowledge base entry in the semantic knowledge base, obtain the keyword vector in the user's question represented by A, identify the entry vector in the semantic knowledge base represented by B, use the cosine similarity algorithm to measure the similarity between the keywords in the user's question and the semantic knowledge base entries, marked as semantic similarity, the cosine similarity algorithm formula is expressed as: The knowledge base entry with the largest semantic similarity is recorded as a semantically related entry; where Ai represents the i-th component of the user question keyword vector A, Bi represents the i-th component of the semantic knowledge base entry vector B, and n represents the dimension of the keyword vector A and the semantic knowledge base entry vector B, that is, the number of elements in the vector; all semantic similarities and their corresponding target knowledge base entries and semantically related entries are recorded as semantic matching results;

[0063] Determine the semantically related items, and parse the problem requirements based on the connotation of the semantically related items and the business rules and data processing logic associated with them; obtain the pre-set data storage structure and processing method, and build the corresponding structured query language or data processing task instruction template; fill the specific parameters of the semantically related items in the questions raised by the user into the structured query language or data processing task instruction template to generate a structured query language or data processing task instruction; send the semantic matching results and the structured query language or data processing task instruction to the reasoning and analysis layer.

[0064] In this application, based on the reasoning and analysis layer, regression analysis in machine learning algorithms is used to analyze the questions raised by users. The specific analysis steps are as follows:

[0065] According to the rule engine, the questions raised by the user are identified and analyzed to determine the independent variable x and the dependent variable y; the linear regression model is selected, and the formula is y=B0+B1x+ε, where B0 and B1 represent regression coefficients and ε represents the error term;

[0066] The regression coefficients are calculated by the least squares method, and B0 and B1 are solved by minimizing the sum of squared errors; the formula is: c represents the number of data points, xt and yt represent the values ​​of the independent variable and dependent variable of the tth data point, respectively. are the means of the independent and dependent variables, respectively.

[0067] In this application, based on the reasoning and analysis layer, cluster analysis in machine learning algorithms is used to analyze the questions raised by users. The specific analysis steps are:

[0068] Set the number of clusters P, where the number of clusters P is determined based on business experience or data analysis results;

[0069] Initialize P cluster centers and randomly select P data points as initial cluster centers;

[0070] The Euclidean distance formula is used to calculate the distance from each data point to the cluster center. The formula is: Among them, xtp represents the value of the t-th data point in the p-th dimension, swp represents the value of the w-th cluster center in the p-th dimension, and m represents the data dimension;

[0071] Assign the data point to the class with the closest cluster center;

[0072] Then recalculate the cluster center and convergence judgment, calculate the average value of the data points in each class as the new cluster center; calculate the difference between the previous cluster center and the current cluster center to get the cluster difference;

[0073] The number of iterations of the current cluster center is counted, and the cluster difference values ​​of all cluster centers and the number of iterations are weighted to obtain the cluster difference value; a cluster difference threshold is set. If the cluster difference value is less than the cluster difference threshold, it means that the cluster center no longer changes or the change range is within a controllable range, and the clustering process converges.

[0074] In this application, based on the number of clusters P, the elbow rule is used to determine the optimal number of clusters, and the specific steps are:

[0075] The sum of squares of each data point and its cluster center is calculated and marked as the sum value SSE. The formula is: Where Q represents the number of cluster centers, Cw represents the w-th cluster center, x represents the data point, and uw represents the center of the w-th cluster;

[0076] Starting from P=2, increase P by one in each round; for each P value, calculate the corresponding aggregation value of each P value; construct a relationship line graph with P value as the horizontal axis and aggregation value as the vertical axis, mark the position of the aggregation value in the relationship line graph as the aggregation point, connect adjacent aggregation points, calculate the slope between adjacent aggregation points, and the formula is g(P)=SSE(P+1)-SSE(P); then calculate the rate of change of its slope Δg(P)=g(P+1)-g(P);

[0077] A change threshold is set. If the change rate is less than the change threshold, the P value corresponding to the change rate is used as the optimal number of clusters. It should be noted that by calculating the sum values ​​under different numbers of clusters and constructing a line graph of the relationship between the sum value and the number of clusters, it is possible to intuitively observe the change of SSE with the increase in the number of clusters. Calculating the slope and change rate between adjacent sum points can help determine the turning point where the SSE change trend changes from steep to slow. The number of clusters corresponding to this turning point is the optimal number of clusters. This method helps to avoid selecting too many or too few clusters in cluster analysis, thereby improving the accuracy and effectiveness of cluster analysis.

[0078] In this application, based on the data collection and integration layer, the frequency of data collection is dynamically adjusted according to the real-time requirements of enterprise data updates, specifically:

[0079] Establish a data update monitoring module to monitor the update timestamp or data change mark of the data in the data source in real time;

[0080] Set a data update time zone, identify the number of times data is updated in the data update time zone and mark it as a number of times; divide the number of times by the duration corresponding to the data update time zone to obtain the data update frequency f;

[0081] Set the normal update frequency range, match the data update frequency with the normal update frequency range, if the data update frequency is within its normal update frequency range, keep the current data collection frequency unchanged;

[0082] Identify the upper limit value T1 and the lower limit value T2 of the normal update frequency range,

[0083] If the data update frequency is greater than or equal to the upper limit of its normal update frequency range, it means that the data is updated frequently. Then the dynamic data acquisition adjustment coefficient is calculated and the formula is expressed as: Use the dynamic data acquisition adjustment coefficient b1 to adjust the data acquisition frequency; if the data update frequency is less than or equal to the lower limit of its normal update frequency range, it means that the data update is not frequent, then calculate and reduce the dynamic data acquisition adjustment coefficient Use the dynamic data acquisition adjustment coefficient b2 to adjust the data collection frequency. It should be noted that the dynamic adjustment mechanism helps to improve the efficiency and accuracy of data collection, especially in scenarios where enterprise data changes frequently and has high requirements for data timeliness. It can ensure that the system always collects data at an appropriate frequency, providing a reliable data basis for subsequent data processing and analysis.

[0084] In the present application, based on the data collection and integration layer, the preprocessing operation includes cleaning, conversion, and format normalization operations; the document form data after the preprocessing operation is stored in a distributed data repository in a unified data storage format and model.

[0085] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art that are not disclosed in this disclosure. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present invention are indicated by the following claims.

[0086] It should be understood that the present invention is not limited to the exact construction that has been described above and shown in the drawings and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A document form processing system based on question-answering reasoning, characterized in that: include: The data collection and integration layer is used to collect the company's document form data and perform pre-processing operations and data storage; The question-answer understanding and conversion layer integrates a natural language processing engine to perform semantic analysis on the questions raised by users and sends the analysis results after semantic analysis to the reasoning and analysis layer; the natural language processing engine includes a word vector model and a syntactic analyzer based on deep learning; Reasoning and analysis layer: design reasoning mechanism based on rule engine and machine learning algorithm. The rule engine predefines business rules and logical relationships. Use machine learning algorithms for pattern recognition, trend prediction and anomaly detection. Machine learning algorithms include regression analysis and cluster analysis. The results obtained by machine learning algorithms are passed to the result presentation and interaction layer. The result presentation and interaction layer uses data visualization technology and report generation tools to present the results in the form of intuitive charts, text reports or interactive data dashboards according to the result type and user needs.

2. The document form processing system based on question-answering reasoning according to claim 1 is characterized in that: Based on the natural language processing engine, the word vector model is used to perform semantic analysis on the questions raised by users. Specifically: Obtain a trained word vector model, use the word vector model to identify the words in the questions raised by the user and convert them into word vectors; adopt a neural network architecture based on the self-attention mechanism, and learn the contextual relationship of the words in the user's questions through a multi-layer self-attention mechanism; Among them, the calculation formula of the self-attention mechanism in the neural network architecture is as follows: For the input sequence X=(x1,x2,…x H ), where H represents the total number of elements in the input sequence X; the query vector Q, key vector K and value vector V are obtained through linear transformation; the formula is: Q = XW Q , K = XW K , V = XW V ;W Q , W K , W V Represents the learnable weight matrix; then calculate its attention score S, the formula is expressed Among them, d k represents the dimension of the key vector K, QK T Matrix multiplication representing the transpose of the query vector and the key vector; Use the Softmax function to normalize the attention score and value vector to obtain the attention weight O, and the formula is O=softmax(S)V.

3. The document form processing system based on question-answering reasoning according to claim 1 is characterized in that: Based on the natural language processing engine, a syntactic analyzer is used to perform syntactic analysis on user questions. The specific steps are as follows: Determine the sentence grammatical structure and identify the non-terminal symbol N in the user's question r With the terminal symbol T v , according to the probabilistic context-free grammar, there is a production rule N r →β, where β represents a string of symbols consisting of non-terminal symbols and terminal symbols, and v represents the index of different terminal symbols in the sentence; For the analysis tree T of any sentence in the user's question, its probability P(T) is obtained by multiplying the probabilities of all production rules, and the formula is expressed as: Among them, Nr→βr represents the production rules in the analysis tree T; r represents the index variable of different symbols in the sentence; Mark the vocabulary vectors and grammatical structures in the user's questions as question analysis results; Build the semantic knowledge base of the enterprise, including enterprise management, financial analysis, academic research and their corresponding professional terms, concept relationships and business rules; match and map the problem analysis results of the user's questions with the entries in the semantic knowledge base, calculate the semantic similarity between the problem analysis results and any knowledge base entry in the semantic knowledge base, obtain the keyword vector in the user's question represented by A, identify the entry vector in the semantic knowledge base represented by B, use the cosine similarity algorithm to measure the similarity between the keywords in the user's question and the semantic knowledge base entries, marked as semantic similarity, the cosine similarity algorithm formula is expressed as: The knowledge base entry with the largest semantic similarity is recorded as a semantically related entry; where Ai represents the i-th component of the user question keyword vector A, Bi represents the i-th component of the semantic knowledge base entry vector B, and n represents the dimension of the keyword vector A and the semantic knowledge base entry vector B, that is, the number of elements in the vector; all semantic similarities and their corresponding knowledge base entries and semantically related entries are recorded as semantic matching results; Determine the semantically related items, and parse the problem requirements based on the connotation of the semantically related items and the business rules and data processing logic associated with them; obtain the pre-set data storage structure and processing method, and build the corresponding structured query language or data processing task instruction template; fill the specific parameters of the semantically related items in the questions raised by the user into the structured query language or data processing task instruction template to generate a structured query language or data processing task instruction; send the semantic matching results and the structured query language or data processing task instruction to the reasoning and analysis layer.

4. The document form processing system based on question-answering reasoning according to claim 1 is characterized in that: Based on the reasoning and analysis layer, the questions raised by users are analyzed using regression analysis in machine learning algorithms. The specific analysis steps are as follows: According to the rule engine, the questions raised by the user are identified and analyzed to determine the independent variable x and the dependent variable y; the linear regression model is selected, and the formula is y=B0+B1x+ε, where B0 and B1 represent regression coefficients and ε represents the error term; The regression coefficients are calculated by the least squares method, and B0 and B1 are solved by minimizing the sum of squared errors; the formula is: c represents the number of data points, xt and yt represent the values ​​of the independent variable and dependent variable of the tth data point, respectively. are the means of the independent and dependent variables, respectively.

5. The document form processing system based on question-answering reasoning according to claim 1 is characterized in that: Based on the reasoning and analysis layer, cluster analysis in machine learning algorithms is used to analyze the questions raised by users. The specific analysis steps are as follows: Set the number of clusters P; Initialize P cluster centers and randomly select P data points as initial cluster centers; The Euclidean distance formula is used to calculate the distance from each data point to the cluster center. The formula is: Among them, xtp represents the value of the t-th data point in the p-th dimension, swp represents the value of the w-th cluster center in the p-th dimension, and m represents the data dimension; Assign the data point to the class with the closest cluster center; Then recalculate the cluster center and convergence judgment, calculate the average value of the data points in each class as the new cluster center; calculate the difference between the previous cluster center and the current cluster center to get the cluster difference; The number of iterations of the current cluster center is counted, and the cluster difference values ​​of all cluster centers and the number of iterations are weighted to obtain the cluster difference value; a cluster difference threshold is set. If the cluster difference value is less than the cluster difference threshold, it means that the cluster center no longer changes or the change range is within a controllable range, and the clustering process converges.

6. The document form processing system based on question-answering reasoning according to claim 5 is characterized in that: Based on the number of clusters P, the elbow rule is used to determine the optimal number of clusters. The specific steps are: Calculate the sum of squares of each data point and its cluster center cluster and mark it as the cluster sum value; Starting from P=2, increase P by one in each round; for each P value, calculate the corresponding cluster sum value for each P value; construct a relationship line graph with P value as the horizontal axis and cluster sum value as the vertical axis, mark the position of the cluster sum value in the relationship line graph as the sum point, connect adjacent sum points, calculate the slope between adjacent sum points, and then calculate the rate of change of the slope; set a change threshold, if the change rate is less than its change threshold, then use the P value corresponding to the change rate as the number of optimal clusters.

7. The document form processing system based on question-answering reasoning according to claim 1 is characterized in that: Based on the data collection and integration layer, the frequency of data collection is dynamically adjusted according to the real-time needs of enterprise data updates, specifically: Establish a data update monitoring module to monitor the update timestamp or data change mark of the data in the data source in real time; Set a data update time zone, identify the number of times data is updated in the data update time zone and mark it as a number of times; divide the number of times by the duration corresponding to the data update time zone to obtain the data update frequency f; Set the normal update frequency range, match the data update frequency with the normal update frequency range, if the data update frequency is within its normal update frequency range, keep the current data collection frequency unchanged; Identify the upper limit value T1 and the lower limit value T2 of the normal update frequency range, If the data update frequency is greater than or equal to the upper limit of its normal update frequency range, it means that the data is updated frequently. Then the dynamic data acquisition adjustment coefficient is calculated and the formula is expressed as: Use the dynamic data acquisition adjustment coefficient b1 to adjust the data acquisition frequency; if the data update frequency is less than or equal to the lower limit of its normal update frequency range, it means that the data update is not frequent, then calculate and reduce the dynamic data acquisition adjustment coefficient Use the dynamic data acquisition adjustment factor b2 to adjust the data acquisition frequency.

8. The document form processing system based on question-answering reasoning according to claim 1 is characterized in that: Based on the data collection and integration layer, the preprocessing operation includes cleaning, conversion, and format normalization operations; the document form data after the preprocessing operation is stored in a distributed data repository in a unified data storage format and model.