Data analysis method, device, computer equipment and storage medium

By analyzing historical texts to generate trigger rules and problem templates, and automatically reviewing application materials, the problem of low manual review efficiency is solved, and a fast and accurate review process is achieved.

CN114218357BActive Publication Date: 2025-08-19PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111537922.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2025-08-19
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

In the prior art, the review of application materials in investment banking business is inefficient and relies on manual operations, which consumes a lot of manpower and material resources and is prone to omissions or calculation errors.

Method used

By analyzing the historical text collection, generating a set of trigger rules, and generating a question template based on the trigger rules, automatically extracting key elements in the target text to match the rules, and generating a question list to quickly review the target text.

Benefits of technology

It realizes automatic review of application materials, improves audit efficiency, saves labor and time costs, reduces omissions and errors, and can quickly generate a list of quality control problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114218357B_ABST
    Figure CN114218357B_ABST
Patent Text Reader

Abstract

The present invention discloses a data analysis method, apparatus, computer equipment, and storage medium, belonging to the field of data processing. The data analysis method analyzes historical texts in a historical text collection to obtain trigger rules, and generates corresponding question templates based on the trigger rules. When a target text to be analyzed is received, key elements in the target text are extracted and matched with the trigger rules in the trigger rule collection to obtain the target trigger rules. The corresponding question template is determined based on the target trigger rules, and a question list is generated based on the question template and the target text, so that the target text can be modified according to the question list, thereby achieving the purpose of quickly reviewing the target text and improving the review efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a data analysis method, apparatus, computer equipment and storage medium. Background Art

[0002] In the investment banking business operation process, in order to control risks, the stock and bond application materials are mainly reviewed by experienced personnel from the perspectives of compliance, debt repayment risk, etc., and quality control or internal review issues (hereinafter collectively referred to as issues) are raised. The project team then modifies, supplements and responds to the application materials based on the issues, and finally forms an application material with comprehensive risk disclosure and complete information disclosure, so as to facilitate submission to the regulatory authorities for review.

[0003] In summary, the application materials are currently mainly reviewed manually based on experience. Due to the complexity of the application materials, it requires a lot of manpower and material resources, is inefficient, and is prone to omissions or calculation errors. Summary of the Invention

[0004] In view of the low efficiency of the current manual review of application materials, a data analysis method, device, computer equipment and storage medium are provided to automatically review application materials with high efficiency.

[0005] To achieve the above object, the present invention provides a data analysis method, comprising:

[0006] Acquire a historical text set, analyze the historical texts in the historical text set to obtain a trigger rule set, and generate corresponding question templates according to each of the trigger rules;

[0007] Obtaining the target text to be analyzed;

[0008] Extracting key elements from the target text, matching the key elements with trigger rules in the trigger rule set, and obtaining target trigger rules that match the key elements;

[0009] A question template corresponding to the target trigger rule is obtained, and a question list is generated according to the question template and the target text.

[0010] Optionally, a historical text set is obtained, the historical texts in the historical text set are analyzed to obtain a trigger rule set, and corresponding question templates are generated according to each trigger rule, including:

[0011] crawling data from a preset website to obtain the historical text collection;

[0012] Classifying the historical texts according to the text types in the historical text set to obtain a rule text subset and a query text subset;

[0013] Analyzing the rule texts in the rule text subset and the query texts in the query text subset to obtain query texts corresponding to the rule texts;

[0014] generating trigger rules according to the rule texts that match the query text, wherein all the trigger rules constitute the trigger rule set;

[0015] The question template is generated according to the trigger rule and the query text.

[0016] Optionally, classifying the historical texts according to the text types in the historical text set to obtain the rule text subset and the query text subset includes:

[0017] Extracting data information of each historical text in the historical text set, wherein the data information includes title, text content and keywords;

[0018] Classifying the historical texts according to the titles to obtain the rule text subset and the query text subset;

[0019] The categories of the historical texts include rule categories and query categories. The rule categories correspond to the rule text subsets, and the query categories correspond to the query text subsets.

[0020] Optionally, analyzing the rule text in the rule text subset and the query text in the query text subset to obtain the query text corresponding to the rule text includes:

[0021] Keywords of the rule texts in the rule text subset are matched with keywords of the query texts in the query text subset to determine query texts corresponding to the respective rule texts.

[0022] Optionally, extracting key elements from the target text, matching the key elements with trigger rules in the trigger rule set, and obtaining target trigger rules matching the key elements includes:

[0023] Extract key elements from the target text;

[0024] According to the key elements, the trigger rule set is parsed using regular expressions to obtain target trigger rules that match the key elements.

[0025] Optionally, the question template includes fill-in items and unfilled items, and each unfilled item corresponds to a category;

[0026] The step of obtaining a question template corresponding to the target trigger rule and generating a question list according to the question template and the target text includes:

[0027] Determining the question template corresponding to the target trigger rule according to the target trigger rule;

[0028] Obtaining identification words of the target text and classifying the identification words;

[0029] adding an identification word that matches the item category of the item to be filled in the question template to the question template to form a question text;

[0030] The question list is generated according to the question text.

[0031] Optionally, each of the question templates is associated with at least one of the rule texts and a query text corresponding to the rule text;

[0032] The question list includes at least one question text, and the question text corresponds to the question template;

[0033] The question list also includes the storage path of the rule text associated with the corresponding question template and the storage path of the inquiry text associated with the corresponding question template.

[0034] To achieve the above object, the present invention provides a data analysis device, comprising:

[0035] An analysis unit, configured to obtain a historical text set, analyze the historical texts in the historical text set, obtain a trigger rule set, and generate a corresponding question template according to each of the trigger rules;

[0036] an acquisition unit, configured to acquire the target text to be analyzed;

[0037] a processing unit, configured to extract key elements from a target text, match the key elements with trigger rules in the trigger rule set, and obtain a target trigger rule that matches the key elements;

[0038] A generating unit is configured to obtain a question template corresponding to the target trigger rule, and generate a question list according to the question template and the target text.

[0039] To achieve the above objectives, the present invention provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0040] To achieve the above object, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.

[0041] The data analysis method, apparatus, computer equipment and storage medium provided by the present invention obtain trigger rules by analyzing historical texts in an acquired historical text set, and generate corresponding question templates based on the trigger rules; when a target text to be analyzed is received, key elements in the target text are extracted, and the key elements are matched with the trigger rules in the trigger rule set to obtain the target trigger rules, and the corresponding question template is determined based on the target trigger rules, thereby generating a question list based on the question template and the target text, so as to facilitate the modification of the target text according to the question list, achieve the purpose of quickly reviewing the target text, and achieve the purpose of improving the review efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A method flow chart of an embodiment of the data analysis method of the present invention;

[0043] Figure 2 A flow chart of a method for generating a question template according to an embodiment of the present invention;

[0044] Figure 3 A flow chart of a method for generating a question list according to a question template and a target text according to an embodiment of the present invention;

[0045] Figure 4 A module diagram of an embodiment of the data analysis device of the present invention;

[0046] Figure 5 This is a module diagram inside the analysis unit of the present invention;

[0047] Figure 6 A block diagram of the processing unit of the present invention;

[0048] Figure 7 FIG. 1 is a schematic diagram of the hardware architecture of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of this application more clear, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0050] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0051] The data analysis method, apparatus, computer equipment, and storage medium provided by the present invention are applicable to the financial sector, such as bank bond issuance. The present invention analyzes historical text from a collection of acquired historical texts to obtain trigger rules, and generates corresponding question templates based on the trigger rules. Upon receiving a target text to be analyzed, the present invention extracts key elements from the target text and matches them with trigger rules from the collection of trigger rules to obtain the target trigger rules. Based on the target trigger rules, the present invention determines a corresponding question template, thereby generating a question list based on the question template and the target text. This allows the target text to be modified based on the question list, enabling rapid review of the target text and improving review efficiency.

[0052] Example 1

[0053] See also Figure 1 , a data analysis method of this embodiment includes the following steps:

[0054] S1. Obtain a historical text set, analyze the historical texts in the historical text set, obtain a trigger rule set, and generate corresponding question templates according to each of the trigger rules.

[0055] In this embodiment, the historical texts in the historical text collection primarily include legal and regulatory texts related to the target text, as well as historical project feedback and inquiry letters related to the target text (e.g., historical bond project feedback and inquiry letters). By analyzing the legal and regulatory texts and historical project feedback and inquiry letters, trigger rules corresponding to the inquiry questions are obtained.

[0056] Further, see Figure 2 As shown, step S1 may include the following steps:

[0057] S11. Crawl data from a preset website to obtain the historical text collection.

[0058] In this embodiment, data from pre-defined websites (e.g., various regulatory websites) can be crawled using a web crawler. A web crawler, also known as a web spider, web robot, or web chaser, is a program or script that automatically crawls the World Wide Web according to specific rules. Web crawlers are also known as ants, automatic indexers, simulation programs, and worms. Web crawlers are implemented in a variety of programming languages, and a large number of plug-ins have been developed for use.

[0059] S12. Classify the historical texts according to the text types in the historical text set to obtain a rule text subset and a query text subset.

[0060] In this embodiment, the document type may include a rule category and an inquiry category. The rule category corresponds to the rule text subset (a collection of legal and regulatory texts), and the inquiry category corresponds to the inquiry text subset (a collection of historical project feedback and inquiry letters). For example, historical project feedback and inquiry letters may include information such as the issuer's industry, bond type, and issuance method.

[0061] Specifically, step S12 may include:

[0062] S121. Extract data information of each historical text in the historical text set, wherein the data information includes title, text content and keywords.

[0063] In this embodiment, the TF-IDF algorithm or the TextRank algorithm may be used to extract data information of the historical text.

[0064] TF-IDF (Term Frequency) measures the number of words that appear most frequently in an article. However, frequently occurring words are not necessarily keywords. For example, common stop words may not be of much significance to the article itself. Therefore, an importance adjustment factor is needed to measure whether a word is common. This weight is the inverse document frequency (IDF), which is inversely proportional to the word's commonness. After obtaining the term frequency (TF) and inverse document frequency (IDF), the two values are multiplied together to obtain a word's TF-IDF value. The more important a word is to the article, the higher its TF-IDF value. Therefore, the top few words in an article are likely to be keywords. The TF-IDF algorithm has the advantages of being simple and fast, and its results are more accurate to actual conditions. However, simply measuring a word's importance based on term frequency is incomplete. Sometimes important words may not appear very often. Moreover, this algorithm fails to reflect word position information, and words that appear at the beginning and those that appear later are considered equally important, which is unreasonable.

[0065] The TextRank algorithm uses a word graph model to find keywords for an article. It's based on the PageRank algorithm, the core algorithm for Google search, which calculates the importance of web pages by looking at the links between them. The PageRank algorithm views the entire internet as a directed graph, with web pages as nodes and links between them as edges. Based on the principle of importance transfer, if web page A contains a link to web page B, the importance ranking of web page B will increase based on the importance of A. When constructing the graph, the TextRank algorithm changes nodes from web pages to words, and links between web pages to co-occurrence relationships between words. In actual processing, a window of a certain length is used, and co-occurrence relationships within the window are considered valid.

[0066] S122. Classify the historical texts according to the titles to obtain the rule text subset and the query text subset.

[0067] In this embodiment, the content of the title is identified, and the title is classified according to the title content. The title categories may include rule categories and query categories.

[0068] S13. Analyze the rule texts in the rule text subset and the query texts in the query text subset to obtain query texts corresponding to the rule texts.

[0069] In this embodiment, NLP (natural language processing) technology may be used to parse the rule texts in the rule text subset and the query texts in the query text subset, thereby determining the query text corresponding to the rule text.

[0070] Specifically, step S13 may include:

[0071] Keywords of the rule texts in the rule text subset are matched with keywords of the query texts in the query text subset to determine query texts corresponding to the respective rule texts.

[0072] In this embodiment, deep learning-based natural language processing technology is used to parse the rule texts in the rule text subset and the query texts in the query text subset. A neural network model is trained with preset samples to generate a matching model. The rule texts are identified based on keywords in the rule texts. The query texts in the query text subset are then fed into the matching model. The matching model extracts keywords, predicts identifiers matching the query texts based on the keywords, and determines the corresponding rule text based on these identifiers.

[0073] S14. Generate trigger rules based on the rule text that matches the query text, and all the trigger rules constitute the trigger rule set.

[0074] For example, if the inquiry is: "Please explain the high proportion of accounts receivable to total assets, and request the reporting accountant to verify the adequacy of the issuer's provision for doubtful accounts receivable and express a clear opinion." The keyword is: "accounts receivable to total assets." Combined with the rule text corresponding to this keyword, the trigger rule is: "Accounts receivable to total assets exceeds 5%."

[0075] S15. Generate the question template according to the trigger rule and the inquiry text.

[0076] It should be noted that each trigger rule corresponds to at least one question template.

[0077] In this embodiment, question information with a high degree of relevance is extracted from all inquiry texts corresponding to the same trigger rule, and a question template is generated based on the question information.

[0078] For example, if the trigger rule is: accounts receivable exceed 5% of total assets at the end of the period, the corresponding question template is: If the issuer's accounts receivable account for 10% of total assets at the end of the period, please explain the reason for the high proportion of accounts receivable to total assets. Please also ask the reporting accountant to verify the adequacy of the issuer's provision for doubtful accounts receivable and provide a clear opinion.

[0079] S2. Obtain the target text to be analyzed.

[0080] In this embodiment, the target text is the application material text to be analyzed. The application material text sent by the user can be obtained through the client.

[0081] S3. Extract key elements from the target text, match the key elements with the trigger rules in the trigger rule set, and obtain target trigger rules that match the key elements.

[0082] Furthermore, step S3 may include:

[0083] S31. Extract key elements from the target text.

[0084] In this embodiment, the TF-IDF algorithm or the TextRank algorithm may be used to extract key elements from the target text.

[0085] S32. Based on the key elements, the trigger rule set is parsed using regular expressions to obtain target trigger rules that match the key elements.

[0086] A regular expression is a text pattern that describes one or more strings to match when searching for text. A regular expression query triggers a rule set based on key elements.

[0087] In this embodiment, the key elements can be calculated according to the preset strategy to obtain the data to be audited. The trigger rule set is queried using a regular expression based on the data to be audited to determine the target trigger rule that matches it. For example, the key elements are: the accounts receivable balance and the total asset balance at the end of the current period. Based on the "accounts receivable balance at the end of the current period" and the "total asset balance", the ratio of the two is calculated (the data to be audited). The trigger rule corresponding to the data to be audited is: the ratio of accounts receivable to total assets at the end of the period exceeds 5%.

[0088] S4. Obtain a question template corresponding to the target trigger rule, and generate a question list based on the question template and the target text.

[0089] It should be noted that: the question template includes fill-in items and items to be filled in, and each item to be filled in corresponds to a category.

[0090] Further, see Figure 3 As shown, step S4 may include the following steps:

[0091] S41. Determine the question template corresponding to the target trigger rule according to the target trigger rule.

[0092] In this embodiment, considering that each trigger rule corresponds to at least one question template, the question template corresponding to the target trigger rule can be determined based on the target trigger rule.

[0093] S42. Obtain identification words of the target text and classify the identification words.

[0094] In this embodiment, the TF-IDF algorithm or the TextRank algorithm may be used to extract the identification words of the target text.

[0095] In this embodiment, the identification words are words used to identify the characteristics of the target text, such as words indicating the time of the target text, the issuing industry involved, the issuing method, the bond type, etc.

[0096] S43. Add an identification word that matches the item category of the item to be filled in the question template to the question template to form a question text.

[0097] S44. Generate the question list according to the question text.

[0098] In this embodiment, each of the question templates is associated with at least one of the rule texts and the inquiry text corresponding to the rule text; the question list includes at least one of the question texts, and the question text corresponds to the question template; the question list also includes the storage path of the rule text associated with the corresponding question template and the storage path of the inquiry text associated with the corresponding question template.

[0099] In this embodiment, the question list includes a storage link of the inquiry case (inquiry text) corresponding to each question text and the laws and regulations (rule text) on which the review is based.

[0100] In this embodiment, the data analysis method obtains trigger rules by analyzing historical texts in the acquired historical text set, and generates corresponding question templates based on the trigger rules; when the target text to be analyzed is received, the key elements in the target text are extracted, and the key elements are matched with the trigger rules in the trigger rule set to obtain the target trigger rules, and the corresponding question template is determined based on the target trigger rules, so as to generate a question list based on the question template and the target text, so as to facilitate the modification of the target text according to the question list, achieve the purpose of quickly reviewing the target text, and achieve the purpose of improving the review efficiency.

[0101] In actual applications, applying data analysis methods to the analysis of bond filing materials for investment business can save manpower and time costs. Based on the target text, a list of questions about quality control issues can be generated in a very short time, greatly improving the efficiency of the quality control core. Data analysis methods can continuously follow up and study regulatory policies and laws and regulations, which can effectively improve analysis and learning efficiency; users can directly view the content of laws and regulations and the content of inquiry cases according to the storage links in the question list, saving the time cost of searching.

[0102] Example 2

[0103] See also Figure 4 A data analysis device 1 of this embodiment includes: an analyzing unit 11, an acquiring unit 12, a processing unit 13 and a generating unit 14.

[0104] The analyzing unit 11 is configured to obtain a historical text set, analyze the historical texts in the historical text set, obtain a trigger rule set, and generate a corresponding question template according to each of the trigger rules.

[0105] In this embodiment, the historical texts in the historical text collection primarily include legal and regulatory texts related to the target text, as well as historical project feedback and inquiry letters related to the target text (e.g., historical bond project feedback and inquiry letters). By analyzing the legal and regulatory texts and historical project feedback and inquiry letters, trigger rules corresponding to the inquiry questions are obtained.

[0106] Further, see Figure 5 As shown, the analysis unit 11 may include: a crawling module 111, a classification module 112, an analysis module 113, a matching module 114 and a generation module 115.

[0107] The crawling module 111 is used to crawl data from a preset website to obtain the historical text collection.

[0108] In this embodiment, data from pre-defined websites (e.g., various regulatory websites) can be crawled using a web crawler. A web crawler, also known as a web spider, web robot, or web chaser, is a program or script that automatically crawls the World Wide Web according to specific rules. Web crawlers are also known as ants, automatic indexers, simulation programs, and worms. Web crawlers are implemented in a variety of programming languages, and a large number of plug-ins have been developed for use.

[0109] The classification module 112 is configured to classify the historical texts according to the text types in the historical text set to obtain a rule text subset and a query text subset.

[0110] In this embodiment, the document type may include a rule category and an inquiry category. The rule category corresponds to the rule text subset (a collection of legal and regulatory texts), and the inquiry category corresponds to the inquiry text subset (a collection of historical project feedback and inquiry letters). For example, historical project feedback and inquiry letters may include information such as the issuer's industry, bond type, and issuance method.

[0111] Specifically, the classification module 112 is used to extract data information of each historical text in the historical text set, where the data information includes title, text content and keywords.

[0112] In this embodiment, the TF-IDF algorithm or the TextRank algorithm may be used to extract data information of the historical text.

[0113] TF-IDF (Term Frequency) measures the number of words that appear most frequently in an article. However, frequently occurring words are not necessarily keywords. For example, common stop words may not be of much significance to the article itself. Therefore, an importance adjustment factor is needed to measure whether a word is common. This weight is the inverse document frequency (IDF), which is inversely proportional to the word's commonness. After obtaining the term frequency (TF) and inverse document frequency (IDF), the two values are multiplied together to obtain a word's TF-IDF value. The more important a word is to the article, the higher its TF-IDF value. Therefore, the top few words in an article are likely to be keywords. The TF-IDF algorithm has the advantages of being simple and fast, and its results are more accurate to actual conditions. However, simply measuring a word's importance based on term frequency is incomplete. Sometimes important words may not appear very often. Moreover, this algorithm fails to reflect word position information, and words that appear at the beginning and those that appear later are considered equally important, which is unreasonable.

[0114] The TextRank algorithm uses a word graph model to find keywords for an article. It's based on the PageRank algorithm, the core algorithm for Google search, which calculates the importance of web pages by looking at the links between them. The PageRank algorithm views the entire internet as a directed graph, with web pages as nodes and links between them as edges. Based on the principle of importance transfer, if web page A contains a link to web page B, the importance ranking of web page B will increase based on the importance of A. When constructing the graph, the TextRank algorithm changes nodes from web pages to words, and links between web pages to co-occurrence relationships between words. In actual processing, a window of a certain length is used, and co-occurrence relationships within the window are considered valid.

[0115] The classification module 112 is further configured to classify the historical text according to the title to obtain the rule text subset and the query text subset.

[0116] In this embodiment, the content of the title is identified, and the title is classified according to the title content. The title categories may include rule categories and query categories.

[0117] The analyzing module 113 is configured to analyze the rule texts in the rule text subset and the query texts in the query text subset to obtain query texts corresponding to the rule texts.

[0118] In this embodiment, NLP (natural language processing) technology may be used to parse the rule texts in the rule text subset and the query texts in the query text subset, thereby determining the query text corresponding to the rule text.

[0119] Specifically, the analysis module 113 is configured to match keywords of the rule texts in the rule text subset with keywords of the query texts in the query text subset to determine query texts corresponding to the respective rule texts.

[0120] In this embodiment, deep learning-based natural language processing technology is used to parse the rule texts in the rule text subset and the query texts in the query text subset. A neural network model is trained with preset samples to generate a matching model. The rule texts are identified based on keywords in the rule texts. The query texts in the query text subset are then fed into the matching model. The matching model extracts keywords, predicts identifiers matching the query texts based on the keywords, and determines the corresponding rule text based on these identifiers.

[0121] The matching module 114 is configured to generate trigger rules according to the rule text that matches the query text, and all the trigger rules constitute the trigger rule set.

[0122] For example, if the inquiry is: "Please explain the high proportion of accounts receivable to total assets, and request the reporting accountant to verify the adequacy of the issuer's provision for doubtful accounts receivable and express a clear opinion." The keyword is: "accounts receivable to total assets." Combined with the rule text corresponding to this keyword, the trigger rule is: "Accounts receivable to total assets exceeds 5%."

[0123] The generating module 115 is configured to generate the question template according to the triggering rule and the inquiry text.

[0124] It should be noted that each trigger rule corresponds to at least one question template.

[0125] In this embodiment, question information with a high degree of relevance is extracted from all inquiry texts corresponding to the same trigger rule, and a question template is generated based on the question information.

[0126] For example, if the trigger rule is: accounts receivable exceed 5% of total assets at the end of the period, the corresponding question template is: If the issuer's accounts receivable account for 10% of total assets at the end of the period, please explain the reason for the high proportion of accounts receivable to total assets. Please also ask the reporting accountant to verify the adequacy of the issuer's provision for doubtful accounts receivable and provide a clear opinion.

[0127] The acquiring unit 12 is configured to acquire the target text to be analyzed.

[0128] In this embodiment, the target text is the application material text to be analyzed. The application material text sent by the user can be obtained through the client.

[0129] The processing unit 13 is configured to extract key elements from the target text, match the key elements with the trigger rules in the trigger rule set, and obtain a target trigger rule that matches the key elements.

[0130] Further, see Figure 6 As shown, the processing unit 13 may include: an extraction module 131 and a parsing module 132 .

[0131] The extraction module 131 is used to extract key elements from the target text.

[0132] In this embodiment, the TF-IDF algorithm or the TextRank algorithm may be used to extract key elements from the target text.

[0133] The parsing module 132 is configured to parse the trigger rule set based on the key elements using regular expressions to obtain target trigger rules that match the key elements.

[0134] A regular expression is a text pattern that describes one or more strings to match when searching for text. A regular expression query triggers a rule set based on key elements.

[0135] In this embodiment, the key elements can be calculated according to the preset strategy to obtain the data to be audited. The trigger rule set is queried using a regular expression based on the data to be audited to determine the target trigger rule that matches it. For example, the key elements are: the accounts receivable balance and the total asset balance at the end of the current period. Based on the "accounts receivable balance at the end of the current period" and the "total asset balance", the ratio of the two is calculated (the data to be audited). The trigger rule corresponding to the data to be audited is: the ratio of accounts receivable to total assets at the end of the period exceeds 5%.

[0136] The generating unit 14 is configured to obtain a question template corresponding to the target trigger rule, and generate a question list according to the question template and the target text.

[0137] It should be noted that: the question template includes fill-in items and items to be filled in, and each item to be filled in corresponds to a category.

[0138] Furthermore, the generating unit 14 is configured to determine the question template corresponding to the target trigger rule according to the target trigger rule.

[0139] In this embodiment, considering that each trigger rule corresponds to at least one question template, the question template corresponding to the target trigger rule can be determined based on the target trigger rule.

[0140] The generating unit 14 is further configured to obtain identification words of the target text and classify the identification words.

[0141] In this embodiment, the TF-IDF algorithm or the TextRank algorithm may be used to extract the identification words of the target text.

[0142] In this embodiment, the identification words are words used to identify the characteristics of the target text, such as words indicating the time of the target text, the issuing industry involved, the issuing method, the bond type, etc.

[0143] The generating unit 14 is further configured to add an identification word that matches the item category of the item to be filled in the question template to the question template to form a question text.

[0144] The generating unit 14 is further configured to generate the question list according to the question text.

[0145] In this embodiment, each of the question templates is associated with at least one of the rule texts and the inquiry text corresponding to the rule text; the question list includes at least one of the question texts, and the question text corresponds to the question template; the question list also includes the storage path of the rule text associated with the corresponding question template and the storage path of the inquiry text associated with the corresponding question template.

[0146] In this embodiment, the question list includes a storage link of the inquiry case (inquiry text) corresponding to each question text and the laws and regulations (rule text) on which the review is based.

[0147] In this embodiment, the data analysis device 1 analyzes the historical texts in the acquired historical text set through the analysis unit 11 to obtain trigger rules, and generates corresponding question templates based on the trigger rules; when the acquisition unit 12 receives the target text to be analyzed, the processing unit 13 is used to extract key elements in the target text, and the key elements are matched with the trigger rules in the trigger rule set to obtain the target trigger rules, and the generation unit 14 is used to determine the corresponding question template based on the target trigger rules, so as to generate a question list based on the question template and the target text, so as to facilitate the modification of the target text according to the question list, achieve the purpose of quickly reviewing the target text, and achieve the purpose of improving the review efficiency.

[0148] In actual applications, when the data analysis device 1 is applied to the analysis of bond declaration materials for investment business, it can save manpower and time costs, and can generate a question list about quality control issues based on the target text in a very short time, greatly improving the efficiency of the quality control core. The data analysis device 1 can continuously follow up and study regulatory policies and laws and regulations, which can effectively improve analysis and learning efficiency; users can directly view the content of laws and regulations and the content of inquiry cases according to the storage links in the question list, saving the time cost of searching.

[0149] Example 3

[0150] To achieve the above-mentioned purpose, the present invention further provides a computer device 2, which includes multiple computer devices 2. The components of the data analysis device 1 of the second embodiment can be dispersed in different computer devices 2. The computer device 2 can be a smart phone, tablet computer, laptop computer, desktop computer, rack server, blade server, tower server or cabinet server (including an independent server or a server cluster composed of multiple servers) that executes the program. The computer device 2 of this embodiment includes at least but is not limited to: a memory 21, a processor 23, a network interface 22 and a data analysis device 1 (refer to Figure 7 ). It should be pointed out that Figure 7The computer device 2 is shown with only components, but it should be understood that implementing all of the components shown is not a requirement and greater or fewer components may alternatively be implemented.

[0151] In this embodiment, the memory 21 includes at least one type of computer-readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 21 can be an internal storage unit of the computer device 2, such as the hard disk or memory of the computer device 2. In other embodiments, the memory 21 can also be an external storage device of the computer device 2, such as a plug-in hard disk equipped on the computer device 2, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, etc. Of course, the memory 21 can also include both the internal storage unit of the computer device 2 and its external storage device. In this embodiment, the memory 21 is generally used to store the operating system and various application software installed on the computer device 2, such as the program code of the data analysis method of Example 1. In addition, the memory 21 can also be used to temporarily store various types of data that have been output or are to be output.

[0152] In some embodiments, the processor 23 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 23 is generally used to control the overall operation of the computer device 2, such as performing control and processing related to data interaction or communication with the computer device 2. In this embodiment, the processor 23 is used to execute program code stored in the memory 21 or process data, such as running the data analysis device 1.

[0153] The network interface 22 may include a wireless network interface or a wired network interface. The network interface 22 is generally used to establish a communication connection between the computer device 2 and other computer devices 2. For example, the network interface 22 is used to connect the computer device 2 to an external terminal via a network, establish a data transmission channel and a communication connection between the computer device 2 and the external terminal, etc. The network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), a 4G network, a 5G network, Bluetooth, or Wi-Fi.

[0154] It should be pointed out that Figure 7 Computer device 2 is shown having only components 21 - 23 , but it should be understood that implementation of all of the components shown is not a requirement, and greater or fewer components may alternatively be implemented.

[0155] In this embodiment, the data analysis device 1 stored in the memory 21 can also be divided into one or more program modules, and the one or more program modules are stored in the memory 21 and executed by one or more processors (processor 23 in this embodiment) to complete the present invention.

[0156] Example 4

[0157] To achieve the above objectives, the present invention also provides a computer-readable storage medium, which includes multiple storage media, such as flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, server, App application store, etc., on which a computer program is stored, and the program realizes the corresponding function when executed by the processor 23. The computer-readable storage medium of this embodiment is used to store the data analysis device 1, and when executed by the processor 23, realizes the data analysis method of embodiment 1.

[0158] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0159] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method.

[0160] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A data analysis method, characterized in that: include: Acquire a historical text set, analyze the historical texts in the historical text set, obtain a trigger rule set, and generate corresponding question templates according to each trigger rule; wherein, the method includes: crawling data from a preset website to obtain the historical text collection; Classifying the historical texts according to text types in the historical text set to obtain a rule text subset and a query text subset; the text types include rule categories and query categories; analyzing the rule texts in the rule text subset and the query texts in the query text subset to obtain query texts corresponding to the rule texts; generating trigger rules according to the rule texts that match the query text, wherein all the trigger rules constitute the trigger rule set; Generating the question template according to the trigger rule and the inquiry text; wherein, highly relevant question information is extracted from all inquiry texts corresponding to the same trigger rule, and the question template is generated based on the question information; Get the target text to be analyzed; Extracting key elements from the target text, matching the key elements with trigger rules in the trigger rule set, and obtaining target trigger rules that match the key elements; A question template corresponding to the target trigger rule is obtained, and a question list is generated according to the question template and the target text.

2. The data analysis method according to claim 1, characterized in that The historical texts are classified according to the text types in the historical text set to obtain a rule text subset and a query text subset, including: Extracting data information of each historical text in the historical text set, wherein the data information includes title, text content and keywords; Classifying the historical texts according to the titles to obtain the rule text subset and the query text subset; The categories of the historical texts include rule categories and query categories. The rule categories correspond to the rule text subsets, and the query categories correspond to the query text subsets.

3. The data analysis method according to claim 2, characterized in that: The analyzing the rule text in the rule text subset and the query text in the query text subset to obtain the query text corresponding to the rule text includes: Keywords of the rule texts in the rule text subset are matched with keywords of the query texts in the query text subset to determine query texts corresponding to the respective rule texts.

4. The data analysis method according to claim 2, wherein: The step of extracting key elements from the target text, matching the key elements with trigger rules in the trigger rule set, and obtaining target trigger rules that match the key elements includes: Extract key elements from the target text; According to the key elements, the trigger rule set is parsed using regular expressions to obtain target trigger rules that match the key elements.

5. The data analysis method according to claim 1, wherein: The question template includes fill-in items and items to be filled in, and each item to be filled in corresponds to a category; The step of obtaining a question template corresponding to the target trigger rule and generating a question list according to the question template and the target text includes: Determining the question template corresponding to the target trigger rule according to the target trigger rule; Obtaining identification words of the target text and classifying the identification words; adding an identification word that matches the item category of the item to be filled in the question template to the question template to form a question text; The question list is generated according to the question text.

6. The data analysis method according to claim 5, characterized in that: Each of the question templates is associated with at least one of the rule texts and a query text corresponding to the rule text; The question list includes at least one question text, and the question text corresponds to the question template; The question list also includes the storage path of the rule text associated with the corresponding question template and the storage path of the inquiry text associated with the corresponding question template.

7. A data analysis device, characterized in that: include: An analysis unit, configured to obtain a historical text set, analyze the historical texts in the historical text set, obtain a trigger rule set, and generate a corresponding question template according to each of the trigger rules; The method includes: crawling data from a preset website to obtain the historical text set; classifying the historical texts according to the text types in the historical text set to obtain a rule text subset and an inquiry text subset; the text types include rule categories and inquiry categories; analyzing the rule texts in the rule text subset and the inquiry texts in the inquiry text subset to obtain inquiry texts corresponding to the rule texts; generating trigger rules based on the rule texts that match the inquiry texts, and all the trigger rules constitute the trigger rule set; generating the question template based on the trigger rules and the inquiry texts; extracting highly relevant question information from all inquiry texts corresponding to the same trigger rule, and generating a question template based on the question information; An acquisition unit, used to acquire the target text to be analyzed; a processing unit, configured to extract key elements from a target text, match the key elements with trigger rules in the trigger rule set, and obtain a target trigger rule that matches the key elements; A generating unit is configured to obtain a question template corresponding to the target trigger rule, and generate a question list according to the question template and the target text.

8. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method and system for automatically extracting and reviewing PCB engineering associated problems

    CN106912162A