An apparatus for AI error correction and traceability explanation of standard terms
Through the RPA data acquisition and AI intelligent error correction module combined with DLF file technology, the problem of improper word use and difficulty in traceability of text is solved, accurate error correction and traceability of text is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510180211.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The prior art is difficult to effectively identify and correct word errors in text, and it is impossible to trace the source of the correct word after correction.
The RPA data acquisition module, DLF file generation module, AI intelligent training module, AI intelligent error correction module, traceability explanation module and combination book generation module are used. Through Internet resources as a knowledge base, traceability archives are carried out according to the GB/T39677-2020 standard, and the AI intelligent error correction module is used for inference and correction, and a dynamic DLF file book is generated to achieve traceability explanation.
It realizes accurate correction and traceability of text content, ensures the standardization and traceability of text, improves user experience, solves the problem of improper text words, and supports localized traceability query.
Abstract
Description
Technical Field
[0001] The present invention relates to the field of AI technology, and in particular to a device for AI error correction and source tracing interpretation of standardized terms. Background Art
[0002] In an era of information explosion, text, as the primary vehicle for information transmission, is crucial for its accuracy and standardization, crucial for its effective communication and understanding. However, due to various factors, such as regional differences, educational levels, and the rise of online slang, text often contains inappropriate word choice and grammatical errors.
[0003] To address this problem, artificial intelligence technology has emerged. Rapid developments in natural language processing (NLP), in particular, have enabled AI systems to deeply understand text content and identify and correct errors in word usage.
[0004] After training on a large number of standard words, the incorrect words can be marked during actual use. However, when combined with sentences, many standard words cannot be accurately identified. After a sentence is used incorrectly and the error is corrected, there is no way to trace the source. Summary of the Invention
[0005] The purpose of the present invention is to provide a device for AI error correction and source tracing of standardized terms to solve the problems raised in the above background technology.
[0006] To achieve the above objectives, the present invention provides the following technical solutions: a device for AI error correction and source interpretation of standardized terms, comprising an RPA data acquisition module, a DLF file generation module, an AI intelligent training module, an AI intelligent error correction module, a source interpretation module, and a combined bound volume generation module.
[0007] The RPA data collection module is used to collect information from the set website and obtain standard terms in the web page based on regular expressions and semantics;
[0008] The DLF file generation module obtains an offline web page file, converts the interface content corresponding to each level in the web page into an OFD file, and assembles the converted OFD files into a DLF file according to the web page hierarchy relationship;
[0009] The AI intelligent training module inputs the collected information into the training application system, extracts the list of "standard words", calls the AI expansion model, and constructs sentences for the specified words in different situations;
[0010] The AI intelligent error correction module is used to infer the user's input statement that needs to be checked or corrected, and return whether there is an error in the statement based on the inference result;
[0011] The source interpretation module is used to parse the data structure returned by the AI intelligent error correction module, render and display the parsed results, mark the source button, and link the parsed results to the original text;
[0012] The combined bound volume generating module combines all linked files in the source tracing interpretation module with the DLF files generated by the DLF file generating module to generate a bound volume.
[0013] Preferably, the RPA data collection module collects information including standardized terms, basic meanings and developments, interpretations, historical backgrounds, corresponding URLs, and offline corresponding web page files.
[0014] Preferably, the DLF file is a compressed package file, which contains a navigation file and several OFD files. The navigation file mainly records the association relationship between files, file entries, storage paths of OFD files, and trigger locations and jump events between OFDs.
[0015] Preferably, when the AI intelligent training module constructs sentences for specified words in different situations, the sentence patterns include simple sentences, parallel sentences, compound sentences, interrogative sentences, exclamatory sentences, declarative imperative sentences, and predicate inversion sentences. When constructing sentences, N sentences of each type are generated according to correct content and incorrect content, and the standardized words, correct sentences, and incorrect sentences are input into the model for sentence training.
[0016] Preferably, the AI intelligent error correction module performs logical reasoning on the input sentences. If it is an erroneous sentence, it marks the erroneous information content: incorrect words, standard normative words, and the location of the error; when the reasoning result returns that there are multiple errors in the sentence, the erroneous content is displayed in the form of a queue, and the erroneous content is directly corrected and then sent to the model for secondary training.
[0017] Preferably, the traceability interpretation module refers to querying the corresponding data information in the normative term storage system based on the normative terminology. The data information includes the normative terminology, basic meaning and development, interpretation, historical background, and website address. After querying the corresponding data information, the page is loaded and a link is established between the page and the original file content.
[0018] The preferred combined bound volume generation module combines all DLF files matched by the traceability interpretation into a bound volume, uses a DLF reader to open the specified DLF file, automatically jumps to the specific traceability location in each DLF, and displays multiple traceability jump directory addresses of each DLF file, which can be switched by clicking.
[0019] Preferably, the combined bound volume generation module supports displaying the specified standardized terms, basic meanings and developments, interpretations, and historical backgrounds on the DLF file using different colors, icons, and styles, and extracting the content to form an outline; when the DLF file displays the traceability content, for the corresponding original text address, first verify whether the original text address exists. If so, directly display the link to the original text on the DLF to achieve the final tracking and tracing of the original address.
[0020] A device for AI error correction and source-tracing interpretation of standardized terms, characterized by:
[0021] Step S1: The RPA data collection module collects information from the specified website, including standardized terms, basic meanings and developments, interpretations, historical backgrounds, corresponding website addresses, and offline corresponding web page files;
[0022] Step S2: The DLF file generation module obtains the offline web page file, converts the interface content corresponding to each level in the web page into an OFD file, and assembles the converted OFD files into a DLF file according to the web page hierarchy relationship;
[0023] Step S3: The AI intelligent training module inputs the collected information into the training application system, calls the AI expansion model, and constructs sentences for the specified words in different situations;
[0024] Step S4: The AI intelligent error correction module infers the sentence input by the user that needs to be checked or corrected, and returns whether there is an error in the sentence based on the inference result;
[0025] Step S5: The source tracing and interpretation module is used to parse the data structure returned by the AI intelligent error correction module and render and display the parsed results;
[0026] Step S6: The combined bound volume generation module combines all linked files in the source tracing interpretation module with the DLF files generated by the DLF file generation module to generate a bound volume.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] (1) This invention creates a new AI-based error correction and source-tracing interpretation device for standardized word usage. When correcting the content of the reading object, it uses Internet resources as its knowledge base and realizes source-tracing comparison;
[0029] (2) This invention archives the traceability object in accordance with the national standard GB / T39677-2020, ensuring its authenticity and integrity in the original form in the current historical period;
[0030] (3) The present invention uses DLF file technology to reassemble archived web pages and generate a dynamic layout file, thereby freezing the original historical appearance of the web page at that time, making the website effect at that time vividly reproduced and able to be preserved for a long time, realizing the retrospective ability to reflect the scene of the web page at that time, greatly improving the user experience;
[0031] (4) The present invention adopts DLF dynamic layout file technology to localize the traceability content, thereby solving the problem of inaccessibility to the external network caused by the separation of the user's internal and external network, and does not affect the user's traceability comparison of the AI error correction of standardized terms;
[0032] (5) The present invention uses an AI intelligent error correction module to reason about sentences. After reasoning, secondary training can be performed on incorrect data, thereby achieving the purpose of continuous learning;
[0033] (6) The present invention converts different web page sources into DLF files and then generates a DLF file binding, so that different web pages can be viewed in the same reader, achieving the purpose of "one-stop search". DETAILED DESCRIPTION
[0034] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0035] A device for AI error correction and source tracing of standardized terms consists of an RPA data acquisition module, a DLF file generation module, an AI intelligent training module, an AI intelligent error correction module, a source tracing and interpretation module, and a combined bound volume generation module.
[0036] The RPA data collection module is used to collect information from the set URLs and obtain the standard words in the web pages based on regular expressions and semantics;
[0037] The DLF file generation module obtains the offline web page file, converts the interface content corresponding to each level in the web page into an OFD file, and assembles the converted OFD files into a DLF file according to the web page hierarchy relationship;
[0038] The AI intelligent training module inputs the collected information into the training application system, extracts the list of "standard words", calls the AI expansion model, and constructs sentences for the specified words in different situations;
[0039] The AI intelligent error correction module is used to infer the user-entered statements that need to be checked or corrected, and returns whether there are errors in the statements based on the inference results;
[0040] The source interpretation module is used to parse the data structure returned by the AI intelligent error correction module, render and display the parsed results, mark the source button, and link the parsed results to the original text;
[0041] The combined bound volume generation module combines all linked files in the source tracing interpretation module with the DLF files generated by the DLF file generation module to generate a bound volume.
[0042] The working steps of a device for AI error correction and source interpretation of standardized terms are as follows:
[0043] Step S1: The RPA data collection module collects information from the specified URL and obtains the standard terms in the web page based on regular expressions and semantics. The collected content includes the standard terms, basic meaning and development, interpretation, historical background, corresponding URL, and offline corresponding web page files;
[0044] Step S2: The DLF file generation module obtains the offline web page file, converts the interface content corresponding to each level in the web page into an OFD file, and assembles the converted OFD files into a DLF file according to the web page hierarchy. The DLF file is a compressed package file that contains a navigation file and several OFD files. The navigation file mainly records the association relationship between files, the entry of the file, the storage path of the OFD file, the triggering position and jump event between OFDs;
[0045] Step S3: The AI intelligent training module inputs the collected information into the training application system, calls the AI expansion model, and constructs sentences for the specified words in different situations. The sentence patterns include simple sentences, parallel sentences, compound sentences, interrogative sentences, exclamatory sentences, declarative imperative sentences, and predicate inversion sentences. When constructing sentences, N sentences of each type are generated according to the correct content and the incorrect content. The standardized words, correct sentences, and incorrect sentences are input into the model for sentence training.
[0046] Step S4: The AI intelligent error correction module infers the user-entered sentence that needs to be checked or corrected, and returns whether the sentence contains errors based on the inference result. For erroneous sentences, the module annotates the error information: incorrect words, standard normative words, and the location of the error. If the inference result returns that the sentence contains multiple errors, the error content is displayed in the form of a queue, and the content determined to be erroneous is directly corrected, and then sent to the model for secondary training.
[0047] Step S5: The source interpretation module is used to parse the data structure returned by the AI intelligent error correction module, render and display the parsed results, and query the corresponding data information in the standardized word storage system based on the standardized word. The data information includes the standardized word, basic meaning and development, interpretation, historical background, and website address. After the corresponding data information is queried, the page is loaded and a link is established between the page and the original file content;
[0048] Step S6: The combined bound volume generation module combines all linked files in the tracing interpretation module with the DLF files generated by the DLF file generation module to generate a bound volume, uses the DLF reader to open the specified DLF file, automatically jumps to the specific tracing location in each DLF, and displays multiple tracing jump directory addresses of each DLF file. Click to switch. The combined bound volume generation module supports displaying the specified standard terms, basic meanings and developments, interpretations, and historical backgrounds on the DLF file with different colors, icons, and styles, and extracts the content to form an outline; when the DLF file displays the tracing content, it first verifies whether the corresponding original text address exists. If it exists, the link to the original text is directly displayed on the DLF to achieve the final tracing of the original address.
[0049] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations that come within the meaning and range of equivalents of the claims be embraced therein.
Claims
1. A device for AI error correction and source interpretation of standardized terms, characterized by: It includes RPA data collection module, DLF file generation module, AI intelligent training module, AI intelligent error correction module, source tracing and interpretation module, and combined bound volume generation module. The RPA data collection module is used to collect information from the set website and obtain standard terms in the web page based on regular expressions and semantics; The DLF file generation module obtains an offline web page file, converts the interface content corresponding to each level in the web page into an OFD file, and assembles the converted OFD files into a DLF file according to the web page hierarchy relationship; The AI intelligent training module inputs the collected information into the training application system, extracts the 'standard word' list, calls the AI expansion model, and constructs sentences for the specified words in different contexts; when the AI intelligent training module constructs sentences for the specified words in different contexts, the sentence patterns include simple sentences, parallel sentences, compound sentences, interrogative sentences, exclamatory sentences, declarative imperative sentences, and predicate inversion sentences. When constructing sentences, N sentences of each type are generated according to correct content and incorrect content, and the standard word, correct sentences, and incorrect sentences are input into the model for sentence training; The AI intelligent error correction module is used to infer the user's input statement that needs to be checked or corrected, and return whether there is an error in the statement based on the inference result; The source interpretation module is used to parse the data structure returned by the AI intelligent error correction module, render and display the parsed results, mark the source button, and link the parsed results to the original text; the source interpretation module refers to querying the corresponding data information in the normative word storage system based on the standard word, the data information includes the standard word, basic meaning and development, interpretation, historical background, and website address. After querying the corresponding data information, the page is loaded and a link is established between the page and the original file content; The combined bound volume generation module combines all linked files in the tracing and interpretation module with the DLF file generated by the DLF file generation module to generate a bound volume. The combined bound volume generation module supports displaying the specified standard terms, basic meanings and developments, interpretations, and historical backgrounds on the DLF file using different colors, icons, and styles, and extracts the content to form an outline. When the DLF file displays the tracing content, it first verifies whether the corresponding original text address exists. If so, the link to the original text is directly displayed on the DLF to achieve the tracing and tracing of the final original address.
2. The device for AI error correction and source interpretation of standardized terms according to claim 1, characterized in that: The RPA data collection module collects information including standardized terms, basic meanings and developments, interpretations, historical backgrounds, corresponding URLs, and offline corresponding web page files.
3. The device for AI error correction and source interpretation of standardized terms according to claim 1, characterized in that: The DLF file is a compressed package file, which contains a navigation file and several OFD files. The navigation file mainly records the association relationship between files, file entries, storage paths of OFD files, and trigger locations and jump events between OFDs.
4. The device for AI error correction and source interpretation of standardized terms according to claim 1, characterized in that: The AI intelligent error correction module performs logical reasoning on the input sentences. If it is an incorrect sentence, it will provide the error information: incorrect words, standard normative words, and the location of the error. When the reasoning result returns that there are multiple errors in the sentence, the error content is displayed in the form of a queue, and the content determined to be incorrect is directly corrected and then sent to the model for secondary training.
5. The device for AI error correction and source interpretation of standardized terms according to claim 1 is characterized by: The combined bound volume generation module combines all DLF files matched by the traceability interpretation into a bound volume, uses a DLF reader to open the specified DLF file, automatically jumps to the specific traceability location in each DLF, and displays multiple traceability jump directory addresses of each DLF file, which can be switched by clicking.
6. The device for AI error correction and source tracing of standardized terms according to claim 1 is characterized by: Step S1: The RPA data collection module collects information from the specified website, including standardized terms, basic meanings and developments, interpretations, historical backgrounds, corresponding website addresses, and offline corresponding web page files; Step S2: The DLF file generation module obtains the offline web page file, converts the interface content corresponding to each level in the web page into an OFD file, and assembles the converted OFD files into a DLF file according to the web page hierarchy relationship; Step S3: The AI intelligent training module inputs the collected information into the training application system, calls the AI expansion model, and constructs sentences for the specified words in different situations; Step S4: The AI intelligent error correction module infers the sentence input by the user that needs to be checked or corrected, and returns whether there is an error in the sentence based on the inference result; Step S5: The source tracing interpretation module is used to parse the data structure returned by the AI intelligent error correction module and render and display the parsed results; Step S6: The combined bound volume generation module combines all linked files in the source tracing interpretation module with the DLF files generated by the DLF file generation module to generate a bound volume.
Citation Information
Patent Citations
Government affair field multi-stage fusion text error correction method based on knowledge graph
CN116502628A
Statistical analysis engine device of unstructured file based on AI semantic understanding
CN119202035A