Knowledge question-answering method and device based on Markdown document knowledge base

By copying Markdown documents in the document knowledge base and using file storage middleware to store audio and video resources, the problem of not being able to display Markdown document audio and video resources in the existing technology is solved, and the effect of users intuitively obtaining document content is achieved, improving user experience.

CN120045517APending Publication Date: 2025-05-27JIANGSU BOZHI SOFTWARE TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510122665.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When saving Markdown documents, the existing document knowledge base only saves text content and cannot display the audio and video resources in the document, which affects the user experience.

Method used

By copying the Markdown document to the training folder and preview folder, and using file storage middleware to store audio and video resources, modify resource links, the display of audio and video resources in the Markdown document is achieved.

Benefits of technology

While ensuring the efficiency of large-scale training, users will be fed back to their users with recommended Markdown documents containing audio and video resources, allowing users to intuitively obtain all the content in the document, improving user experience and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045517A_ABST
    Figure CN120045517A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge question and answer method and device based on a Markdown document knowledge base, and the method comprises the steps: copying each preview Markdown document to a same-name sub-folder, storing video and audio resources associated with each preview Markdown document into a file storage middleware, and modifying a resource link of each video and audio resource in each preview Markdown document; the method comprises the following steps: outputting a recommended Markdown document according to a user input problem by using a large language model obtained by training a candidate Markdown document; generating a document acquisition link according to the selected recommended Markdown document name, and acquiring a recommended Markdown document according to the document acquisition link; according to the method, the audio and video resources are obtained in the file storage middleware according to the resource link in the recommended Markdown document, and the recommended Markdown document containing the audio and video resources is displayed to the user, so that the user satisfaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent question answering, and in particular, to a knowledge question answering method and device based on a Markdown (a lightweight markup language) document knowledge base. Background Art

[0002] With the rapid development of information technology, especially the wide application of artificial intelligence, machine learning, and natural language processing technologies, the application scope of large models is becoming more and more extensive. Currently, in order to address the new challenges of data security and privacy protection brought about by the wide application of large models, more and more enterprises and organizations have started to perform private deployment of large models.

[0003] In a private deployment environment, it is usually necessary to build a document knowledge base based on documents in formats such as Markdown, text, Portable Document Format (PDF), PowerPoint (PPT), Microsoft Office Word (word), and Microsoft Office Excel (excel), and through the large model, train and learn on its own based on the existing document knowledge base without an Internet connection.

[0004] However, since only pure text information is stored in text documents, and documents such as pdf, ppt, word, and excel can store audio-visual resources embedded in the file and then store them as a single file, while the text content and audio-visual resources in Markdown documents are usually stored separately. Therefore, the existing document knowledge base usually only saves the text content in Markdown documents, resulting in the inability to display the audio-visual resources in Markdown documents to users, thus affecting the user experience. Summary of the Invention

[0005] The present invention provides a knowledge question answering method and device based on a Markdown document knowledge base, which can, while ensuring the training efficiency of the large model, feedback a recommended Markdown document containing audio-visual resources to the user, enabling the user to intuitively obtain all the content in the recommended Markdown document, and improving the user experience and satisfaction.

[0006] In a first aspect, an embodiment of the present invention provides a knowledge question answering method based on a Markdown document knowledge base, and the method includes:

[0007] Copy each Markdown document to the training folder, and correct the format of the training Markdown documents in the training folder to obtain candidate Markdown documents;

[0008] Copy each Markdown document to a subfolder with the same name as each Markdown document, and store the subfolders with the same name in the preview folder;

[0009] Store the audio-visual resources associated with each preview Markdown document in the preview folder in a pre-constructed file storage middleware, and modify the resource links of each audio-visual resource in each preview Markdown document according to the name of each audio-visual resource in the file storage middleware and the storage path of the file storage middleware;

[0010] Output a text-form answer and a preset number of recommended Markdown documents with the highest relevance to the user input question according to the large model trained by using the candidate Markdown documents;

[0011] In response to the preview operation triggered by the user for the target recommended Markdown document, generate a document acquisition link according to the name of the target recommended Markdown document, and obtain the target recommended Markdown document in the preview folder according to the document acquisition link;

[0012] Obtain each audio-visual resource in the file storage middleware according to the resource links recorded in the target recommended Markdown document, and feedback the text content and audio-visual resources in the target recommended Markdown document to the user for preview.

[0013] Optionally, before copying each Markdown document to the training folder, it further includes: checking and correcting the document name, title, and resource link prefix of each Markdown document.

[0014] Optionally, check and correct the document name, title, and resource link prefix of each Markdown document, including: identifying the title identification characters in the Markdown document, and adding a space after the title identification characters when the characters after the title identification characters are not spaces; checking the resource link prefix in the Markdown document, and correcting the resource link prefix in the Markdown document to the standard resource link prefix when the resource link prefix in the Markdown document is different from the standard resource link prefix; checking whether the document name of the Markdown document contains spaces, and replacing the spaces in the document name with underscores when the document name of the Markdown document contains spaces; checking whether the Markdown document has the same name as another Markdown document, and sending a prompt message to the user indicating that the currently processed Markdown document has the same name as another Markdown document when the Markdown document has the same name as another Markdown document.

[0015] Optionally, correct the format of the training Markdown documents in the training folder to obtain candidate Markdown documents, including: deleting the directory file keywords and redundant keywords in the training Markdown documents; identifying the resource insertion strings and resource format strings in the training Markdown documents, and deleting the strings starting with the resource insertion strings and ending with the resource format strings; identifying the first-level headings and second-level headings in the training Markdown documents, and adding a title identification character before the first character of the training Markdown document when no first-level heading is detected in the training Markdown document, so as to use the first character of the training Markdown document as the first-level heading.

[0016] Optionally, the audio-visual resources include pictures, videos, and audios; storing the audio-visual resources associated with each preview Markdown document in the preview folder into a pre-constructed file storage middleware, including: obtaining the pre-defined picture storage volume, video storage volume, and audio storage volume in the file storage middleware, and storing the pictures associated with each preview Markdown document into the picture storage volume, storing the videos associated with each preview Markdown document into the video storage volume, and storing the audios associated with each preview Markdown document into the audio storage volume; modifying the resource links of each audio-visual resource in each preview Markdown document according to the names of each audio-visual resource in the file storage middleware and the storage path of the file storage middleware, including: determining the storage path of each storage volume according to the storage path of the file storage middleware and the names of each storage volume; modifying the resource links of each audio-visual resource in each preview Markdown document according to the names of each audio-visual resource in the file storage middleware and the storage path of each storage volume.

[0017] Optionally, a large model trained by using candidate Markdown documents outputs a preset number of recommended Markdown documents with the highest relevance to the user input question according to the user input question, including: representing the user input question and each Markdown document as vectors by the large model, and determining the similarity between the user input question and each Markdown document according to the similarity between the vectors; using the preset number of Markdown documents with the highest similarity to the user input question as the recommended Markdown documents by the large model.

[0018] Optionally, obtaining each audio-visual resource in the file storage middleware according to each resource link recorded in the target recommended Markdown document, including: generating anti-theft resource links corresponding to each resource link respectively according to each resource link recorded in the target recommended Markdown document and the anti-theft link corresponding to the target recommended Markdown document; obtaining each audio-visual resource in the file storage middleware according to each anti-theft resource link.

[0019] In a second aspect, an embodiment of the present invention further provides a knowledge Q&A device based on a Markdown document knowledge base, and the device includes:

[0020] A candidate document obtaining module, configured to copy each Markdown document to a training folder, and correct the format of the training Markdown documents in the training folder to obtain candidate Markdown documents;

[0021] A preview document acquisition module, configured to copy each Markdown document to a subfolder with the same name corresponding to each Markdown document respectively, and store each subfolder with the same name in a preview folder;

[0022] A video and audio resource transfer and storage module, configured to store the video and audio resources associated with each preview Markdown document in the preview folder in a pre-constructed file storage middleware, and modify the resource links of each video and audio resource in each preview Markdown document according to the names of each video and audio resource in the file storage middleware and the storage path of the file storage middleware;

[0023] A knowledge Q&A module, configured to output text-form answers and a preset number of recommended Markdown documents with the highest relevance to the user input question according to the user input question by using a large model trained with candidate Markdown documents;

[0024] A document acquisition module, configured to generate a document acquisition link according to the name of a target recommended Markdown document in response to a preview operation triggered by the user for the target recommended Markdown document, and acquire the target recommended Markdown document in the preview folder according to the document acquisition link;

[0025] A document preview module, configured to acquire each video and audio resource in the file storage middleware according to each resource link recorded in the target recommended Markdown document, and feed back the text content and video and audio resources in the target recommended Markdown document to the user for preview together.

[0026] In a third aspect, an embodiment of the present invention further provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the knowledge Q&A method based on a Markdown document knowledge base provided in any embodiment of the present invention.

[0027] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores computer instructions for causing a processor to implement the knowledge Q&A method based on a Markdown document knowledge base provided in any embodiment of the present invention when executed.

[0028] The technical solution provided by the embodiment of the present invention copies each Markdown document to the training folder and the preview folder, trains a large model through the training Markdown documents in the training folder, and feeds back a target recommended Markdown document containing audio-visual resources to the user through the preview Markdown documents in the preview folder. It avoids the situation that the existing document knowledge base only saves the text content in the Markdown document, resulting in the inability to display the audio-visual resources in the Markdown document to the user. It can feed back the recommended Markdown document containing audio-visual resources to the user while ensuring the training efficiency of the large model, enabling the user to intuitively obtain all the content in the recommended Markdown document, improving the user experience and satisfaction, and enhancing the practicality and interactivity of the Markdown document knowledge base.

[0029] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0031] Figure 1 is a flowchart of a knowledge Q&A method based on a Markdown document knowledge base according to Embodiment 1 of the present invention;

[0032] Figure 2 is an effect diagram of knowledge Q&A by a large model according to a Markdown document knowledge base provided by an embodiment of the present invention;

[0033] Figure 3 is a flowchart of another knowledge Q&A method based on a Markdown document knowledge base according to Embodiment 2 of the present invention;

[0034] Figure 4 is a schematic structural diagram of a knowledge Q&A device based on a Markdown document knowledge base according to Embodiment 3 of the present invention;

[0035] Figure 5 is a schematic structural diagram of an electronic device provided by Embodiment 4 of the present invention. Detailed Description of the Embodiments

[0036] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0038] Embodiment 1

[0039] Figure 1 is a flowchart of a knowledge Q&A method based on a Markdown document knowledge base according to Embodiment 1 of the present invention. This embodiment is applicable to the situation of performing knowledge Q&A through a large model based on a Markdown document knowledge base. This method can be executed by a knowledge Q&A device based on a Markdown document knowledge base. The knowledge Q&A device based on a Markdown document knowledge base can be implemented in the form of hardware and / or software, and the knowledge Q&A device based on a Markdown document knowledge base can be configured in an electronic device such as a computer.

[0040] As Figure 1 shown, a knowledge Q&A method based on a Markdown document knowledge base disclosed in this embodiment includes:

[0041] S110. Copy each Markdown document to the training folder, and correct the format of the training Markdown documents in the training folder to obtain candidate Markdown documents.

[0042] In this embodiment, the training Markdown documents can be the Markdown documents copied to the training folder. The Markdown documents can include text content and resource links of audio-visual resources. The audio-visual resources can include pictures, videos, and audios, etc.

[0043] In this step, specifically, Markdown documents can be obtained according to the suffix names of each pre-collected backup document, and non-Markdown backup documents can be converted into Markdown documents. Exemplarily, assuming there are backup documents with suffix names of ".md", ".txt", ".pdf", ".ppt", ".word", and ".excel", the backup documents with suffix names of ".txt", ".pdf", ".ppt", ".word", and ".excel" can be converted into Markdown documents with the suffix of ".md".

[0044] Then, each Markdown document can be copied to the training folder, and at least one of the title, document name, and resource link prefix of each training Markdown document can be checked and corrected to obtain candidate Markdown documents.

[0045] S120: Copy each Markdown document to a subfolder with the same name corresponding to each Markdown document, and store each subfolder with the same name in the preview folder.

[0046] In this step, specifically, a subfolder with the same name corresponding to each Markdown document can be constructed according to the document name of each Markdown document. Then, each Markdown document can be stored in the subfolder with the same name corresponding to it. Among them, the document name of the Markdown document does not contain the suffix name of the Markdown document.

[0047] S130: Store the audio-visual resources associated with each preview Markdown document in the pre-constructed file storage middleware, and modify the resource links of each audio-visual resource in each preview Markdown document according to the name of each audio-visual resource in the file storage middleware and the storage path of the file storage middleware.

[0048] In this step, specifically, the audio-visual resources associated with each preview Markdown document can be renamed, and the renamed audio-visual resources can be stored in the file storage middleware. Then, the storage path of the file storage middleware can be concatenated with the name of the audio-visual resource in the file storage middleware to obtain an updated resource link corresponding to each audio-visual resource, and the resource link of each audio-visual resource in each preview Markdown document can be modified to the corresponding updated resource link. Optionally, since the names of the same audio-visual resource in different preview Markdown documents may be different, duplicate removal processing can be performed on the same audio-visual resource with different names before the renaming process of each audio-visual resource.

[0049] The advantage of this setting is that, compared with the prior art that stores audio-visual resources through local relative paths, the technical solution of this embodiment can ensure the stable loading of audio-visual resources by centrally storing each audio-visual resource in the file storage middleware, avoiding the situation of audio-visual resource loss or loading failure.

[0050] It should be noted that the candidate Markdown documents in the training folder, the preview Markdown documents in the preview folder, and the audio-visual resources in the file storage middleware together constitute the Markdown document knowledge base of the present invention.

[0051] S140. Using the large model trained with the candidate Markdown documents, output a text-form answer to the user input question and a preset number of recommended Markdown documents with the highest correlation degree to the user input question.

[0052] In this embodiment, the large model can be a machine learning model including a retriever and a generator.

[0053] In this step, specifically, the retriever can calculate the similarity between the user input question and each candidate Markdown document, and determine the correlation degree between each candidate Markdown document and the user input question according to the similarity between the user input question and each candidate Markdown document. In practical applications, the higher the similarity between a Markdown document and the user input question, the higher the correlation degree between it and the user input question. The retriever queries information related to the user input question in the candidate Markdown document with the highest similarity, and the generator generates a text-form answer to the user input question according to the information related to the user input question queried by the retriever.

[0054] S150. In response to the preview operation triggered by the user for the target recommended Markdown document, generate a document acquisition link according to the name of the target recommended Markdown document, and obtain the target recommended Markdown document in the preview folder according to the document acquisition link.

[0055] In this step, specifically, a document acquisition link can be generated according to the name of the target recommended Markdown document and the storage path of the preview folder, and the target recommended Markdown document can be obtained in the subfolder with the same name as the target recommended Markdown document according to the document acquisition link.

[0056] S160. According to each resource link recorded in the target recommended Markdown document, obtain each audio-visual resource in the file storage middleware, and jointly feedback the text content and audio-visual resources in the target recommended Markdown document to the user for preview.

[0057] Figure 2 This is an effect diagram of a knowledge Q&A through a large model based on a Markdown document knowledge base according to an embodiment of the present invention. As Figure 2 shown, the knowledge Q&A interface provided to the user includes a user Q&A panel and a knowledge question panel. Among them, the user Q&A panel includes a "New Conversation" button and a list of historical conversation records. The user can start a new conversation by clicking the "New Conversation" button and obtain the historical conversation record by clicking the historical conversation in the list of historical conversation records.

[0058] The knowledge question panel includes a question input box, a question display box, a text answer display box, and a recommended Markdown document display box. In practical applications, in response to the question input by the user in the question input box, the user input question can be displayed in the question display box. Then, through the large model, the text form answer can be displayed in the text answer display box according to the user input question, and the recommended Markdown document can be displayed in the recommended Markdown document display box. After that, in response to the user's click operation on the "Preview" button displayed on the right side of the second recommended Markdown document, the document acquisition link of the second recommended Markdown document can be obtained according to the document name of the second recommended Markdown document through reverse proxy technology, and the second recommended Markdown document can be obtained according to the document acquisition link. Finally, the second recommended Markdown document containing pictures, playable videos, and playable audios can be fed back to the user in the form of a pop-up dialog box.

[0059] The technical solution of this embodiment is to obtain candidate Markdown documents by copying each Markdown document into a training folder and correcting the format of the training Markdown documents in the training folder; copying each Markdown document into a subfolder with the same name corresponding to each Markdown document, and storing each subfolder with the same name into a preview folder; storing the audio and video resources associated with each preview Markdown document in the preview folder into a pre-built file storage middleware, and modifying the resource link of each audio and video resource in each preview Markdown document according to the name of each audio and video resource in the file storage middleware and the storage path of the file storage middleware; outputting a text answer according to the user input question and a preset number of recommended Markdown documents with the highest correlation with the user input question by using the large model trained with the candidate Markdown documents. file; in response to a preview operation triggered by a user for a target recommended Markdown document, a document acquisition link is generated according to the name of the target recommended Markdown document, and the target recommended Markdown document is acquired in the preview folder according to the document acquisition link; according to each resource link recorded in the target recommended Markdown document, each audio and video resource is acquired in the file storage middleware, and the text content and audio and video resources in the target recommended Markdown document are fed back to the user for preview by a technical means, which solves the problem that the existing document knowledge base only saves the text content in the Markdown document, resulting in the inability to display the audio and video resources in the Markdown document to the user. While ensuring the training efficiency of the large model, the recommended Markdown document containing audio and video resources can be fed back to the user, so that the user can intuitively obtain all the content in the recommended Markdown document, thereby improving the user's usage experience and satisfaction.

[0060] Embodiment 2

[0061] Figure 3 It is a flowchart of another knowledge question and answer method based on the Markdown document knowledge base provided according to the second embodiment of the present invention. This embodiment is a further optimization and expansion based on the above embodiments, and can be combined with various optional technical solutions in the above implementation methods.

[0062] like Figure 3 As shown, this embodiment discloses a knowledge question answering method based on a Markdown document knowledge base, including:

[0063] S210. After checking and correcting the document name, title, and resource link prefix of each Markdown document, copy each Markdown document to the training folder and correct the format of the training Markdown documents in the training folder to obtain candidate Markdown documents.

[0064] In this step, specifically, it is possible to check whether the formats of the titles, resource link prefixes, and document names of each Markdown document are correct, and whether the document names of each Markdown document are repeated, and correct the Markdown documents with incorrect formats and repeated document names.

[0065] Optionally, checking and correcting the document name, title, and resource link prefix of each Markdown document includes: identifying the title identification characters in the Markdown document, and adding a space after the title identification characters when the characters after the title identification characters are not spaces; checking the resource link prefix in the Markdown document, and correcting the resource link prefix in the Markdown document to the standard resource link prefix when the resource link prefix in the Markdown document is different from the standard resource link prefix; checking whether the document name of the Markdown document contains spaces, and replacing the spaces in the document name with underscores when the document name of the Markdown document contains spaces; checking whether the Markdown document is a duplicate Markdown document, and sending a prompt message that the currently processed Markdown document is a duplicate Markdown document to the user when the Markdown document is a duplicate Markdown document.

[0066] Among them, the title identification characters can be used to reflect the level of the title. For example, the title identification character of a first-level title is "#", the title identification character of a second-level title is "##", the title identification character of a third-level title is "", and the title identification character of a fourth-level title is "#". The resource link prefix can be "!![]", and the reference to audio-visual resources usually appears in the form of "!![]audio-visual resource.audio-visual resource suffix name".

[0067] Exemplarily, assuming that a first-level title is "#First Title", it can be considered that there is a lack of a space between "#" and "First Title", and at this time, a space can be added between "#" and "First Title". Assuming that the reference to a picture is "!!【】picture.png", it can be considered that the resource link prefix of the picture is incorrect, and at this time, the reference to the picture can be corrected to "!![]picture.png".

[0068] The advantages of such settings are as follows. By checking and correcting the document name, title, and resource link prefix of each Markdown document, it is possible to avoid situations where the large model cannot correctly recognize the overall structure of the Markdown document due to inconsistent Markdown document formats, and at the same time cannot accurately recognize the logical relationship between the title and the body text, thus affecting the training effect of the large model. Secondly, by sending a prompt message to the user that the currently processed Markdown document is a duplicate Markdown document, allowing the user to modify the document name of the duplicate Markdown document, it is possible to ensure that the document names of the preview Markdown documents stored in the preview folder are unique and conflict-free, avoiding the situation where the target recommended Markdown document cannot be accurately obtained based on the document name of the target recommended Markdown document in the future, and improving user satisfaction.

[0069] Optionally, formatting the training Markdown documents in the training folder to obtain candidate Markdown documents includes: deleting the table of contents file keywords and redundant keywords in the training Markdown documents; identifying the resource insertion strings and resource format strings in the training Markdown documents, and deleting the strings starting with the resource insertion strings and ending with the resource format strings; identifying the first-level headings and second-level headings in the training Markdown documents, and when no first-level headings are detected in the training Markdown documents, adding a title identification character before the first line of characters in the training Markdown document to use the first line of characters in the training Markdown document as the first-level heading.

[0070] Among them, the table of contents file keyword can be "[TOC]". Redundant keywords can be used to guide users to pay attention to audio-visual resources. There can be multiple redundant keywords, such as "as shown in the figure", "as shown in the following figure", "as shown in the figure below", "as shown in the picture", "as shown in the audio", and "as shown in the video", etc. The resource insertion string can be "![](", and the resource format string can be ".audio-visual resource suffix)".

[0071] Specifically, when no second-level headings are detected in the training Markdown documents and the number of training Markdown documents without second-level headings reaches the preset document quantity, a list of training Markdown documents without second-level headings can be generated and sent to the user to prompt the user that the content of the above training Markdown documents without second-level headings is less and needs to be excluded or corrected.

[0072] Exemplarily, strings for referencing image resources that start with "![(" and end with ".png", ".jpg", ".jpeg", ".bmp", ".tif", and ".tiff" etc. can be deleted. Strings for referencing video resources that start with "![(" and end with ".wmv", ".mpg", ".mpeg", ".mov", ".mp4", and ".avi" etc. can be deleted. Strings for referencing audio resources that start with "![(" and end with ".mp3", ".wav", ".aac", ".flac", ".ogg", and ".m4a" etc. can be deleted.

[0073] The advantage of this setting is that by deleting the content related to audio-visual resources in the training Markdown document, the invalid information in the training Markdown document can be reduced, and the accuracy of the large model's understanding of the document structure and content during training can be improved.

[0074] S220. Copy each Markdown document to a subfolder with the same name corresponding to each Markdown document, and store each subfolder with the same name in the preview folder.

[0075] Among them, the preview folder is located on the server where the large model server is located.

[0076] S230. Obtain the pre-defined image storage volume, video storage volume, and audio storage volume in the file storage middleware, store the images associated with each preview Markdown document in the image storage volume, store the videos associated with each preview Markdown document in the video storage volume, and store the audio associated with each preview Markdown document in the audio storage volume.

[0077] Among them, the file storage middleware is located on the server where the large model server is located.

[0078] In this step, specifically, the images, videos, and audio associated with each preview Markdown document can be renamed, and the renamed images, videos, and audio can be stored in the corresponding storage volumes. Exemplarily, assuming the name of the image associated with the preview Markdown document is "aaa.png", then after modifying the above image name to "123e4567e89b12d3a45642661417400.png", the renamed image can be stored in the image storage volume.

[0079] S240. Determine the storage paths of each storage volume according to the storage path of the file storage middleware and the names of each storage volume, and modify the resource links of each audio-visual resource in each preview Markdown document according to the names of each audio-visual resource in the file storage middleware and the storage paths of each storage volume.

[0080] In this step, specifically, the storage path of the file storage middleware and the name of the storage volume can be concatenated to obtain the storage paths of each storage volume. Then, the storage path of the storage volume and the name of the audio-visual resource in the file storage middleware can be concatenated to obtain the updated resource link, and the resource link of each audio-visual resource in each preview Markdown document can be corrected to the corresponding updated resource link.

[0081] Exemplarily, assuming that the reference of the picture is "![](aaa.png)", the name of this picture in the file storage middleware is "123e4567e89b12d3a45642661417400.png", and the storage path of the picture storage volume is " / api-gateway-server / api-ai-center", then the resource link of this picture in the preview Markdown document can be modified to " / api-gateway-server / api-ai-center / 123e4567e89b12d3a45642661417400.png".

[0082] It should be noted that since the file storage middleware and the large model server are on the same server, the IP address prefix of the Internet Protocol (IP) does not need to be added to the replaced resource link.

[0083] S250. Output a text-form answer and a preset number of recommended Markdown documents with the highest relevance to the user input question according to the large model trained by using the candidate Markdown documents.

[0084] Optionally, the large model trained by using the candidate Markdown documents outputs a preset number of recommended Markdown documents with the highest relevance to the user input question, including: representing the user input question and each Markdown document as vectors by the large model, and determining the similarity between the user input question and each Markdown document according to the similarity between the vectors; using the large model to take the preset number of Markdown documents with the highest similarity to the user input question as the recommended Markdown documents.

[0085] Specifically, the similarity between each Markdown document vector and the user input question vector can be calculated through a large model, and a preset number of Markdown documents with the highest similarity are used as the recommended Markdown documents.

[0086] The advantage of this setting is that by using a preset number of Markdown documents with the highest similarity to the user input question as the recommended Markdown documents, it is convenient for users to quickly obtain the required document materials, thus improving the user experience.

[0087] S260. In response to the preview operation triggered by the user for the target recommended Markdown document, generate a document acquisition link according to the name of the target recommended Markdown document, and obtain the target recommended Markdown document in the preview folder according to the document acquisition link.

[0088] S270. Generate anti-theft resource links corresponding to each resource link according to each resource link recorded in the target recommended Markdown document and the anti-theft link corresponding to the target recommended Markdown document.

[0089] In this step, specifically, the anti-theft link corresponding to the target recommended Markdown document can be obtained in the preset anti-theft link list according to the document name of the target recommended Markdown document, and the anti-theft resource links corresponding to each resource link are generated according to the resource link and the anti-theft link. Among them, the preset anti-theft link list can be used to store dynamically generated anti-theft links. Exemplarily, assuming the anti-theft link is 9e107d9d372bb6829173b4cc4a0d826a, then “ / api-gateway-server / api-ai-center / 123e4567e89b12d3a45642661417400.png” can be modified to “ / api-gateway-server / api-ai-center / 123e4567e89b12d3a45642661417400.png?token=9e107d9d372bb6829173b4cc4a0d826a”.

[0090] Optionally, the expiration time of each anti-theft link can be set in the file storage middleware, and the anti-theft resource link is generated according to the resource link, the anti-theft link and the timestamp to avoid the situation that the audio-visual resources are stolen for a long time due to the loss of the anti-theft link.

[0091] S280. Obtain each audio-visual resource in the file storage middleware according to each anti-theft resource link, and feedback the text content and audio-visual resources in the target recommended Markdown document to the user for preview.

[0092] In this step, specifically, the target recommended Markdown document can be opened and rendered in the form of a dialog box to display the target recommended Markdown document containing audio-visual resources to the user.

[0093] The technical solution of this embodiment can avoid the situation that the Markdown document cannot be correctly loaded due to format errors in the Markdown document, thus affecting the training effect of the large model by checking and correcting the document name, title, and resource link prefix of each Markdown document. Secondly, by using the preset number of Markdown documents with the highest similarity to the user input question as the recommended Markdown documents, it is convenient for users to quickly obtain the required document materials, thereby improving the user experience.

[0094] Embodiment III

[0095] Figure 4 is a schematic structural diagram of a knowledge Q&A device based on a Markdown document knowledge base according to Embodiment III of the present invention, as Figure 4 shown. The device includes: a candidate document acquisition module 41, a preview document acquisition module 42, an audio-visual resource transfer and storage module 43, a knowledge Q&A module 44, a document acquisition module 45, and a document preview module 46, where:

[0096] The candidate document acquisition module 41 is used to copy each Markdown document to the training folder and correct the format of the training Markdown documents in the training folder to obtain candidate Markdown documents;

[0097] The preview document acquisition module 42 is used to copy each Markdown document to a subfolder with the same name corresponding to each Markdown document and store the subfolders with the same name in the preview folder;

[0098] The audio-visual resource transfer and storage module 43 is used to store the audio-visual resources associated with each preview Markdown document in the preview folder in a pre-constructed file storage middleware, and modify the resource links of each audio-visual resource in each preview Markdown document according to the name of each audio-visual resource in the file storage middleware and the storage path of the file storage middleware;

[0099] The knowledge Q&A module 44 is used to output a text answer and a preset number of recommended Markdown documents with the highest relevance to the user input question according to the large model trained using the candidate Markdown documents;

[0100] The document acquisition module 45 is used to respond to the preview operation triggered by the user for the target recommended Markdown document, generate a document acquisition link according to the name of the target recommended Markdown document, and acquire the target recommended Markdown document in the preview folder according to the document acquisition link;

[0101] The document preview module 46 is used to acquire each audio-visual resource in the file storage middleware according to each resource link recorded in the target recommended Markdown document, and feedback the text content and audio-visual resources in the target recommended Markdown document to the user for preview.

[0102] In the technical solution of this embodiment, through the mutual cooperation of the candidate document acquisition module, the preview document acquisition module, the audio-visual resource transfer and storage module, the knowledge Q&A module, the document acquisition module and the document preview module, the problem that the existing document knowledge base only saves the text content in the Markdown document, resulting in the inability to display the audio-visual resources in the Markdown document to the user is solved. While ensuring the training efficiency of the large model, the recommended Markdown document containing audio-visual resources can be fed back to the user, so that the user can intuitively obtain all the content in the recommended Markdown document, improving the user experience and satisfaction.

[0103] Optionally, the device further includes a document preprocessing module, which is used to: check and correct the document name, title and resource link prefix of each Markdown document.

[0104] Optionally, the document preprocessing module is specifically used to: identify the title identification characters in the Markdown document, and add a space after the title identification characters in the current Markdown document when the characters after the title identification characters are not spaces; check the resource link prefix in the Markdown document, and correct the resource link prefix in the Markdown document to the standard resource link prefix when the resource link prefix in the Markdown document is different from the standard resource link prefix; check whether the document name of the Markdown document contains spaces, and replace the spaces in the document name with underscores when the document name of the Markdown document contains spaces; check whether the Markdown document is a duplicate Markdown document, and send a prompt message indicating that the currently processed Markdown document is a duplicate Markdown document to the user when the Markdown document is a duplicate Markdown document.

[0105] Optionally, the candidate document acquisition module 41 is specifically configured to: delete the directory file keywords and redundant keywords in the training Markdown document; identify the resource insertion strings and resource format strings in the training Markdown document, and delete the strings starting with the resource insertion strings and ending with the resource format strings; identify the first-level headings and second-level headings in the training Markdown document, and when the first-level headings are not detected in the training Markdown document, add a heading identification character before the first-line characters of the training Markdown document to use the first-line characters of the training Markdown document as the first-level headings.

[0106] Optionally, when storing the audio-visual resources associated with each preview Markdown document in the preview folder into the pre-constructed file storage middleware, the audio-visual resource transfer module 43 is specifically configured to: obtain the pre-defined picture storage volume, video storage volume, and audio storage volume in the file storage middleware, and store the pictures associated with each preview Markdown document into the picture storage volume, store the videos associated with each preview Markdown document into the video storage volume, and store the audio associated with each preview Markdown document into the audio storage volume.

[0107] Optionally, when modifying the resource links of each audio-visual resource in each preview Markdown document according to the names of each audio-visual resource in the file storage middleware and the storage path of the file storage middleware, the audio-visual resource transfer module 43 is specifically configured to: determine the storage paths of each storage volume according to the storage path of the file storage middleware and the names of each storage volume; modify the resource links of each audio-visual resource in each preview Markdown document according to the names of each audio-visual resource in the file storage middleware and the storage paths of each storage volume.

[0108] Optionally, the knowledge Q&A module 44 is specifically configured to: represent the user input question and each Markdown document as vectors through a large model, and determine the similarity between the user input question and each Markdown document according to the similarity between the vectors; use the preset number of Markdown documents with the highest similarity to the user input question as the recommended Markdown documents through the large model.

[0109] Optionally, the document preview module 46 is specifically configured to: generate anti-theft resource links corresponding to each resource link according to each resource link recorded in the target recommended Markdown document and the anti-theft link corresponding to the target recommended Markdown document; obtain each audio-visual resource in the file storage middleware according to each anti-theft resource link.

[0110] The knowledge Q&A device based on the Markdown document knowledge base provided by the embodiments of the present invention can execute the knowledge Q&A method based on the Markdown document knowledge base provided by any embodiment 1 of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. The content not described in detail in this embodiment can be referred to the description in any method embodiment of this application.

[0111] Embodiment 4

[0112] Figure 5 Fig. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention.

[0113] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0114] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0115] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the knowledge Q&A method based on the Markdown document knowledge base.

[0116] In some embodiments, the knowledge - based question - answering method based on a Markdown document knowledge base can be implemented as a computer program tangibly embodied in a computer - readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the knowledge - based question - answering method based on a Markdown document knowledge base described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the knowledge - based question - answering method based on a Markdown document knowledge base in any other suitable way (e.g., by means of firmware).

[0117] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field - programmable gate arrays (FPGA), application - specific integrated circuits (ASIC), application - specific standard products (ASSP), systems - on - a - chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special - purpose or general - purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0118] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand - alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0119] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0120] In order to provide interaction with a client user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the client user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the client user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the client user; for example, the feedback provided to the client user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the client user can be received in any form (including acoustic input, voice input, or tactile input).

[0121] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a client user computer having a graphical client user interface or a web browser through which the client user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of the communication network include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0122] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0123] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0124] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A knowledge question answering method based on a Markdown document knowledge base, characterized in that: The method comprises: Copy each Markdown document to the training folder, and correct the format of the training Markdown document in the training folder to obtain a candidate Markdown document; Copy each Markdown document to a subfolder with the same name corresponding to each Markdown document, and store each subfolder with the same name in the preview folder; The audio and video resources associated with each preview Markdown document in the preview folder are stored in the pre-built file storage middleware, and the resource link of each audio and video resource in each preview Markdown document is modified according to the name of each audio and video resource in the file storage middleware and the storage path of the file storage middleware; The large model trained with candidate Markdown documents outputs a text answer based on the user's input question, as well as a preset number of recommended Markdown documents that are most relevant to the user's input question. In response to a preview operation triggered by a user on a target recommended Markdown document, a document acquisition link is generated according to the name of the target recommended Markdown document, and the target recommended Markdown document is acquired in a preview folder according to the document acquisition link; According to the resource links recorded in the target recommended Markdown document, the audio and video resources are obtained in the file storage middleware, and the text content and audio and video resources in the target recommended Markdown document are fed back to the user for preview.

2. The method according to claim 1, characterized in that Before copying each Markdown document into the training folder, also include: Check and correct the document name, title, and resource link prefix of each Markdown document.

3. The method according to claim 2, characterized in that Check and correct the document name, title, and resource link prefix of each Markdown document, including: Identify the title identification characters in the Markdown document, and add a space after the title identification characters of the current Markdown document if the characters after the title identification characters are not spaces; Check the resource link prefix in the Markdown document, and if the resource link prefix in the Markdown document is different from the standard resource link prefix, correct the resource link prefix in the Markdown document to the standard resource link prefix; Checks whether the document name of a Markdown document contains spaces, and replaces the spaces in the document name with underscores if the document name of a Markdown document contains spaces; Check whether a Markdown document has a duplicate name, and if so, send a prompt message to the user that the Markdown document currently being processed is a duplicate name Markdown document.

4. The method according to claim 1, characterized in that: Correct the format of the training Markdown document in the training folder to obtain the candidate Markdown document, including: Delete the directory file keywords and redundant keywords in the training Markdown document; Identify resource insertion strings and resource format strings in the training Markdown document, and delete the strings that start with the resource insertion string and end with the resource format string; First-level titles and second-level titles in a training Markdown document are identified, and when a first-level title is not detected in the training Markdown document, title identification characters are added before the first line of characters in the training Markdown document, so that the first line of characters in the training Markdown document is used as a first-level title.

5. The method according to claim 1, characterized in that The audio-visual resources include pictures, videos and audios; Store the audio and video resources associated with each preview Markdown document in the preview folder into the pre-built file storage middleware, including: Obtain the image storage volume, video storage volume, and audio storage volume pre-defined in the file storage middleware, and store the images associated with each preview Markdown document in the image storage volume, store the videos associated with each preview Markdown document in the video storage volume, and store the audios associated with each preview Markdown document in the audio storage volume; According to the name of each audio and video resource in the file storage middleware and the storage path of the file storage middleware, modify the resource link of each audio and video resource in each preview Markdown document, including: Determine the storage path of each storage volume according to the storage path of the file storage middleware and the name of each storage volume; According to the name of each audio and video resource in the file storage middleware and the storage path of each storage volume, the resource link of each audio and video resource in each preview Markdown document is modified.

6. The method according to claim 1, characterized in that The large model trained with the candidate Markdown documents outputs a preset number of recommended Markdown documents with the highest relevance to the user input question, including: The user input question and each Markdown document are represented as vectors through the large model, and the similarity between the user input question and each Markdown document is determined based on the similarity between the vectors; The large model uses a preset number of Markdown documents with the highest similarity to the user input question as recommended Markdown documents.

7. The method according to claim 1, characterized in that According to the resource links recorded in the target recommended Markdown document, various audio and video resources are obtained in the file storage middleware, including: According to each resource link recorded in the target recommended Markdown document and the anti-hotlink corresponding to the target recommended Markdown document, generate an anti-theft resource link corresponding to each resource link; According to each anti-theft resource link, obtain each audio and video resource in the file storage middleware.

8. A knowledge question-answering device based on a Markdown document knowledge base, characterized in that: The device comprises: A candidate document acquisition module is used to copy each Markdown document into a training folder and perform format correction on the training Markdown document in the training folder to obtain a candidate Markdown document; A preview document acquisition module is used to copy each Markdown document into a subfolder with the same name corresponding to each Markdown document, and store each subfolder with the same name into a preview folder; The audio-visual resource transfer module is used to store the audio-visual resources associated with each preview Markdown document in the preview folder into the pre-built file storage middleware, and modify the resource link of each audio-visual resource in each preview Markdown document according to the name of each audio-visual resource in the file storage middleware and the storage path of the file storage middleware; The knowledge question answering module is used to output a text answer based on the user's input question by using a large model trained with candidate Markdown documents, as well as a preset number of recommended Markdown documents with the highest relevance to the user's input question; A document acquisition module, for generating a document acquisition link according to the name of the target recommended Markdown document in response to a preview operation triggered by a user for the target recommended Markdown document, and acquiring the target recommended Markdown document in the preview folder according to the document acquisition link; The document preview module is used to obtain various audio and video resources in the file storage middleware according to the resource links recorded in the target recommended Markdown document, and feed back the text content and audio and video resources in the target recommended Markdown document to the user for preview.

9. An electronic device, characterized in that: The electronic device comprises: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the knowledge question and answer method based on the Markdown document knowledge base as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the knowledge question and answer method based on a Markdown document knowledge base according to any one of claims 1 to 7 when executed.

Citation Information

Cited By

  • Knowledge grading extraction method for scientific and technical literature in coal industry

    CN120216699A

  • File interaction method and device, equipment, medium and product

    CN122331805A