AI original text accurate reading device with knowledge base containing DLF dynamic layout files
Through the knowledge base AI original text reading device of DLF dynamic layout files, the problem of low relevance of explanatory texts in large-model question and answer is solved, the traceability display of dynamic websites and accurate reading of multiple sources are realized, and the user experience and the accuracy of explanatory texts are improved.
Patent Information
- Application Number
- CN202510185915.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-02-20
AI Technical Summary
In the large-model question-answering process, the explanation text has a low correlation with the input question, cannot be displayed as the original version, and cannot associate multiple sources for inference and presentation, which affects the accuracy of the explanation text and the judgment of the question and answer content.
It uses a knowledge base AI original text precision reading device containing DLF dynamic layout files. Through the data offline collection module, DLF file generation module, content extraction knowledge base storage module, AI reasoning module, tracing and interpretation module and combination comparison module, it realizes the tracing and combination reading of web page content, improves the accuracy of text interpretation and the associated reasoning of multiple sources.
It realizes the traceability display of dynamic websites or application systems, improves the user experience, supports accurate AI original text tracing of various file types, enhances reading accuracy and accuracy of text interpretation, and can display the joint reasoning display of multiple sources as the original.
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular is an AI original text accurate reading device containing a knowledge base of a DLF dynamic layout file. Background Art
[0002] In the common large-scale question-and-answer process, the answer generation process and the explanation text generation process are independent of each other. There may be problems such as low correlation between the explanation text and the input question, and the explanation text does not contain the information in the question, which affects the accuracy of the explanation text. There is no reference and comparison method for the content of the question and answer, and it is impossible to determine whether the content of the answer is accurate. It is impossible to display the original version as it is. When there are multiple sources, there is no way to associate the original sources and jointly perform reasoning and display. Summary of the Invention
[0003] The purpose of the present invention is to provide a knowledge base AI original text accurate reading device containing DLF dynamic layout files to solve the problems raised in the above background technology.
[0004] To achieve the above objectives, the present invention provides the following technical solutions: a knowledge base AI original text accurate reading device containing DLF dynamic layout files, including a data offline acquisition module, a DLF file generation module, a content extraction knowledge base storage module, an AI reasoning module, a source tracing interpretation module, and a combination comparison module.
[0005] The data collection offline module sets the corresponding collection URL, analyzes the web page structure, and generates the web page offline according to the hierarchical relationship of the web page;
[0006] The DLF file generation module converts each offline web page file into an OFD file, and assembles multiple web pages into a dynamic layout DLF file according to the hierarchical relationship. Then, according to national standards, the web pages are archived into an OFD file, which is the source of traceability;
[0007] The content extraction knowledge base storage module is used to extract content information from each individual web page and generate a corresponding file queue based on the extracted content information. The generated file queue is pushed to the inference model to form a knowledge base containing DLF dynamic layout files.
[0008] The AI reasoning module performs logical reasoning on the questions inputted into the content extraction knowledge base storage module within the knowledge base and returns the reasoning results;
[0009] The source interpretation module matches the associated DLF file based on the data returned by the AI reasoning module and performs combined reading;
[0010] The combined comparison module can select multiple page areas for simultaneous comparison and viewing based on different DLF files.
[0011] Preferably, the data collection offline module records the jump relationship between different web pages during the offline process.
[0012] Preferably, the DLF file generation module is to obtain the offline corresponding web page file in the data acquisition offline module, convert each web page into an OFD file, and assemble multiple web pages into a dynamic layout DLF file according to the relationship between the levels.
[0013] Preferably, the content extraction knowledge base storage module first stores the offline DLF file in the landing system, and the specific steps of extracting data are as follows: Step S1: first parse the DLF file format and extract the content information of each individual OFD web page archive file; Step S2: for each page, use MD5 to perform HASH operation on it to form a file queue, where the key is the hash value and the value is the content of the page; Step S3: form a self-describing json data for the DLF file; Step S4: push the file queue generated after parsing each page in the DLF file into the AI reasoning model to form a knowledge base containing DLF dynamic layout files, where the files in the knowledge base come from a single or multiple combinations or merged readable files in the file system.
[0014] Preferably, the AI reasoning module obtains the data pushed in the content extraction module, and outputs the returned data in a certain format, wherein the file array combination of the upper page and lower page of the file is matched to the returned fileld source in the file system JSON; and the obtained fileld combination is passed into the AI reasoning model, and the final result is obtained according to the actual situation, and the source of the knowledge is confirmed to be accurate by reverse reasoning confirmation. Finally, the data is organized and converted into a structured array.
[0015] Preferably, the returned data includes the HASH value of the OFD web page archive file and the specific source content.
[0016] Preferably, the tracing interpretation module obtains the data returned by the AI reasoning module, obtains the fileld field information, and matches the associated DLF file. The specific process when performing combined reading is: extract multiple DLF files; combined reading, and parsing the default jump to load it onto the file page; loading the current page and marking the source, and continuing to preload the fileld sub-level surface.
[0017] Preferably, the combined comparison module creates multiple file display areas from top to bottom during the reading and display process for different DLF files. The left menu uses partitions as page numbers, and the right side switches to display files; on the reader window, a file with a hierarchical structure is constructed, and the file is marked with the source. When clicking to switch, different DLFs and specific pages can be switched. When viewing, several page areas can be selected for comparison and viewing at the same time. The DLF file includes converting files in other formats into DLF files.
[0018] In another aspect, the present invention provides a method for accurately reading original text in a knowledge base AI format containing a DLF dynamic layout file, comprising the following steps:
[0019] Step S1: Set the collection URL and analyze the web page structure, offline the web page according to the hierarchical relationship of the web page, and record the jump relationship between the web pages in real time;
[0020] Step S2: Obtain the above offline web pages, convert each web page into an OFD file, and assemble them into a dynamic layout DLF file according to the hierarchical relationship;
[0021] Step S3: extract the content information of each web page in the DLF file, and perform operations on each page to form a file queue. The file queue generated after parsing each page in the DLF is pushed into the inference model;
[0022] Step S4: Use the reasoning model to perform logical reasoning within the knowledge base for the input question, and return the reasoning results in a certain format.
[0023] Step S5: Match the returned data in the file system to the corresponding upper and lower pages of the content. The matching file combination is then passed into the inference model to obtain the inference structure and perform reverse inference to confirm.
[0024] Step S6: Obtain the matching DLF files, read them together, and mark the source;
[0025] Step S7: Use a reader to display and read multiple DLF files, and compare and view multiple page areas at the same time.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] (1) The present invention creatively defines the source of the content of a dynamic Internet website or application system. According to the national standard GB / T39677-2020, the source is defined to a specific historical moment, and the DLF dynamic layout technology is used to perfectly display the dynamic moment of the application system or application website at this historical moment, thereby achieving traceability and improving the user experience.
[0028] (2) The present invention is aimed at accurate AI original text reading of the knowledge base, not only for DLF files, but also for other files, including various readable files in the knowledge base that are single or multiple combinations or merges from the file system, thereby achieving accurate AI original text tracing of various types of files.
[0029] (3) The present invention can perform combined reading on DLF files and simultaneously match the sources of multiple question and answer results in the knowledge base for comparison, not only including the source of the file and the location of the file, but also combining the upper and lower files for comparison reading, thereby greatly improving the reading accuracy.
[0030] (4) The present invention may be displayed as is. When there are multiple sources, the sources of the related original texts may be combined for inference and display. DETAILED DESCRIPTION
[0031] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. Example
[0032] The present invention provides a knowledge base AI original text precision reading device containing DLF dynamic layout files. The dynamic layout files are referred to as DLF. DLF exists in the form of a file package, which contains a group of OFD files and their associated and displayed relationships to form an assembly, so as to realize the collection of layout files and complex dynamic operations. The device of the present invention includes a data offline acquisition module, a DLF file generation module, a content extraction knowledge base storage module, an AI reasoning module, a source tracing interpretation module, and a combination comparison module.
[0033] The data collection offline module gathers different web page contents according to the corresponding collection URL, analyzes the web page structure, and offline web pages according to the hierarchical relationship of the web pages. During the offline process, the jump relationship between different web pages is recorded. The jump relationship between web pages can be expressed as clicking on a content or part of the content on web page A, and then jumping to another web page B related to the content, or clicking on a content or part of the content on web page A, and then jumping to web page C related to the content, or clicking on the same content on web page A will jump to multiple related web pages such as web pages B and C, thereby realizing the jump between web pages.
[0034] After the data collection is completed, the specific process of generating the DLF file in the DLF file generation module is to obtain the offline web pages in the data collection offline module, and then convert each offline web page file into an OFD file, and assemble multiple web pages into a dynamic layout DLF file according to the hierarchical relationship. According to the national standard GB / T39677-2020, the web pages are archived as OFD files, and the OFD file is used as the source of traceability when tracing back later.
[0035] The content extraction knowledge base storage module obtains the web pages collected by the data collection module, extracts the content information of each individual web page, and generates a corresponding file queue for the extracted content information. The generated file queue is pushed into the inference model to form a knowledge base containing DLF dynamic layout files. In the process of forming the knowledge base containing DLF dynamic layout files, the content extraction knowledge base storage module first stores the offline DLF files in the landing system, and then the specific steps of extracting data are as follows: Step S1: first parse the DLF file format and extract the content information of each individual OFD web page archive file; Step S2: For each page, use MD5 to perform HASH operation on it to form a file queue, where the key is the hash value and the value is the content of the page; Step S3: Generate a self-describing json data for the DLF file, where the hierarchical structure between files is reflected in the json data; Step S4: Push the file queue generated after parsing each page in the DLF file to the AI inference model to form a knowledge base containing DLF dynamic layout files, where the files in the knowledge base may come from a single file, a combined file, or a merged file in the file system.
[0036] The AI reasoning module performs logical reasoning within the knowledge base based on the input questions, obtains the data pushed by the content extraction knowledge base storage module, and outputs the returned data in a certain format. For the returned fileld source, the file array combination of the upper page and lower page of the file is matched in the file system JSON; and the obtained fileld combination is passed into the AI reasoning model. According to the actual situation, the final result is obtained, and the source of the knowledge is confirmed to be accurate through reverse reasoning. Finally, the data is organized and converted into a structured array. The returned data includes the HASH value of the OFD webpage archive file and the specific source content. For example, the returned data format includes:
[0037] {
[0038] 'answer':'',
[0039] 'source': [
[0041] {'fileId':'HASH value of the web page file'
[0042] 'content': 'Source specific source content
[0043] 'sort': sort number
[0044] } ]
[0046] }.
[0047] The traceability interpretation module matches the associated DLF file based on the data returned by the AI reasoning module. The OFD file archived in the DLF file generation module is used as the traceability source, and then the matched DLF files are combined and read. The specific process of combined reading is: extract multiple DLF files; combine and read, and parse the default jump and load it onto the file page; load the current page and mark the source. The source is the OFD file archived in the DLF file generation, and continue to preload the fileld sub-level surface, so that it can be more streamlined when reading multiple levels, greatly improving the reading experience.
[0048] The combined comparison module selects multiple page areas for simultaneous comparison and viewing for different DLF files. During the reading display process, multiple file display areas are created from top to bottom. The left menu uses partitions as page numbers, and the right side switches to display files. In the reader window, a file with a hierarchical structure is constructed, and the source is marked on the file. When clicking to switch, you can switch between different DLFs and specific pages. When viewing, you can select several page areas for comparison and simultaneous viewing.
[0049] In another aspect, the present invention provides a method for accurately reading original text in a knowledge base AI format containing a DLF dynamic layout file, comprising the following steps:
[0050] Step S1: Set the collection URL and analyze the web page structure, offline the web page according to the hierarchical relationship of the web page, and record the jump relationship between the web pages in real time;
[0051] Step S2: Obtain the above offline web pages, convert each web page into an OFD file, and assemble them into a dynamic layout DLF file according to the hierarchical relationship;
[0052] Step S3: extract the content information of each web page in the DLF file, and perform operations on each page to form a file queue, forming a knowledge base containing DLF files, and push the file queue generated after parsing each page in the DLF into the inference model;
[0053] Step S4: Perform logical reasoning within the knowledge base using the reasoning model for the input question and return the result of the reasoning;
[0054] Step S5: Match the returned data in the file system. If the corresponding content's upper and lower pages are matched, it means that the matching not only matches the currently associated content, but also the upper and lower pages of the content. Then, the matched file combination is continuously passed into the inference model to obtain the inference structure and perform reverse inference confirmation.
[0055] Step S6: Obtain the matching DLF files, perform combined reading, and mark the source, that is, trace the source through the OFD file and mark it in the DLF file;
[0056] Step S7: Use a reader to display and read multiple DLF files, and compare and view multiple page areas at the same time.
[0057] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments are to be considered in all respects as illustrative and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations that come within the meaning and range of equivalents of the claims be embraced therein.
Claims
1. A precise AI original text reading device with a knowledge base containing DLF dynamic layout files, characterized by: Including data offline acquisition module, DLF file generation module, content extraction knowledge base storage module, AI reasoning module, traceability explanation module, combination comparison module, The data offline collection module sets the corresponding collection URL, analyzes the web page structure, and collects the web pages offline according to the hierarchical relationship of the web pages; The DLF file generation module converts each offline web page file into an OFD file, assembles multiple web pages into a dynamic layout DLF file according to the hierarchical relationship, and archives the web pages into an OFD file, which is the source of traceability; The content extraction knowledge base storage module is used to extract content information from each individual web page, and generate a corresponding file queue for the extracted content information, and push the generated file queue to the inference model to form a knowledge base containing DLF dynamic layout files; the content extraction knowledge base storage module first stores the offline DLF file in the landing system, wherein the specific steps of extracting data are as follows: Step S1: first parse the DLF file format and extract the content information from each individual OFD web page archive file; Step S2: for each page, perform a HASH operation on it using the MD5 method to form a file queue, wherein the key is the hash value and the value is the content of the page; Step S3: generate a self-describing json data for the DLF file; Step S4: push the file queue generated after parsing each page in the DLF file to the AI inference model to form a knowledge base containing DLF dynamic layout files, wherein the files in the knowledge base come from a single or multiple combinations or merged readable files of the file system; The AI reasoning module performs logical reasoning on the questions input into the content extraction knowledge base storage module within the knowledge base and returns the reasoning results; the AI reasoning module obtains the data pushed in the content extraction module and outputs the returned data, wherein the fileld source returned is matched to the file array combination of the upper page and lower page of the file in the file system JSON; and the obtained fileld combination is passed into the AI reasoning model. According to the actual situation, the final result obtained is confirmed by reverse reasoning to confirm whether the source of the knowledge is accurate, and finally the data is organized and converted into a structured array; The source interpretation module matches the associated DLF file based on the data returned by the AI reasoning module and performs combined reading. The source interpretation module obtains the data returned by the AI reasoning module, obtains the fileld field information, and matches the associated DLF file. The specific process of combined reading is as follows: extract multiple DLF files; combine reading, parse the default jump and load it onto the file page; load the current page and mark the source, and continue to preload the fileld sub-level surface; The combined comparison module can select multiple page areas for simultaneous comparison and viewing based on different DLF files.
2. The precise AI original text reading device containing a knowledge base of DLF dynamic layout files according to claim 1, characterized in that: The data collection offline module records the jump relationship between different web pages during the offline process.
3. The precise AI original text reading device containing a knowledge base of DLF dynamic layout files according to claim 1, characterized in that: The DLF file generation module is to obtain the offline corresponding web page files in the data acquisition offline module, convert each web page into an OFD file, and assemble multiple web pages into a dynamic layout DLF file according to the hierarchical relationship.
4. The precise AI original text reading device containing a knowledge base of DLF dynamic layout files according to claim 1, characterized in that: The returned data includes the HASH value of the OFD webpage archive file and the specific source content.
5. The precise AI original text reading device containing a knowledge base of DLF dynamic layout files according to claim 1 is characterized in that: The combined comparison module creates multiple file display areas from top to bottom during the reading and display process for different DLF files. The left menu uses partitions as page numbers, and the right side switches to display files. On the reader window, a file with a hierarchical structure is constructed, and the source of the file is marked on the file. When clicking to switch, different DLFs and specific pages can be switched. When viewing, several page areas can be selected for comparison and viewing at the same time.
6. The precise AI original text reading device containing a knowledge base of DLF dynamic layout files according to claim 1 is characterized in that: The DLF file includes files in other formats converted into a DLF file.
7. The precise AI original text reading device containing a knowledge base of DLF dynamic layout files according to claim 1, characterized in that: The following steps are involved: Step S1: Set the collection URL and analyze the web page structure, offline the web page according to the hierarchical relationship of the web page, and record the jump relationship between the web pages in real time; Step S2: Obtain the above offline web pages, convert each web page into an OFD file, and assemble them into a dynamic layout DLF file according to the hierarchical relationship; Step S3: extract the content information of each web page in the DLF file, and perform operations on each page to form a file queue. The file queue generated after parsing each page in the DLF is pushed into the inference model; Step S4: Perform logical reasoning within the knowledge base using the reasoning model for the input question and return the result of the reasoning; Step S5: Match the returned data in the file system to the corresponding upper and lower pages of the content. The matching file combination is then passed into the inference model to obtain the inference structure and perform reverse inference to confirm. Step S6: Obtain the matching DLF files, read them together, and mark the source; Step S7: Use a reader to display and read multiple DLF files, and compare and view multiple page areas at the same time.
Citation Information
Patent Citations
Dynamic layout file DLF generation device
CN117910438A
Method for accurately distinguishing and comparing knowledge base and task object based on combined reading
CN118485149A