A method for constructing a question-answering system based on a large language model and a question-answering system.

By constructing a question-answering system based on a large language model, this paper solves the problems of poor understanding of user input and inability to intelligently review and update data in the port and shipping industry. It realizes intelligent data review and update, and improves data accuracy and the efficiency of the question-answering system.

CN118964587BActive Publication Date: 2025-10-28NEZHA SMART TECHNOLOGY (SHANGHAI) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411455126.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2025-10-28
Estimated Expiration
2044-10-18

AI Technical Summary

Technical Problem

Existing question-and-answer systems in the port and shipping sector have poor ability to understand user-input text and cannot achieve intelligent review and updating of data entering the database, resulting in untimely knowledge updates and difficulty in ensuring accuracy.

Method used

A question-answering system is built based on a large language model. By acquiring and preprocessing port and shipping data, review rules and modules are generated. Multiple review large language models are used to retrieve, enhance, generate, and review the data entering the database. Combined with port and shipping industry knowledge and standards, the accuracy of the data entering the database is ensured. The data is entered into the database after passing the review of multiple models.

Benefits of technology

It enables intelligent review and updating of incoming data, improving data accuracy and efficiency, reducing manual review costs, and enhancing the question-and-answer system's ability to understand user input and the correctness of output answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118964587B_ABST
    Figure CN118964587B_ABST
Patent Text Reader

Abstract

This invention discloses a method for constructing a question-answering system based on a large language model and a question-answering system. The method includes: preprocessing port and shipping data to generate data for storage; generating review rules using a large language model based on port and shipping industry knowledge and standards, with the review rules operating through a combination of retrieval-enhanced generation and prompt words; generating a review module using the large language model based on the review rules, the review module including review knowledge text and preset review prompt words, and obtaining review retrieval information based on retrieval-enhanced generation; and reviewing the data for storage using multiple large language models based on the review module and the review retrieval information, and entering the data into the port and shipping database when all multiple large language models have passed the review. This invention achieves intelligent review / update of data for storage, at least solving the problem that existing databases cannot achieve intelligent review / update of data for storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent question answering technology, specifically to a method for constructing a question answering system based on a large language model and a question answering system. Background Technology

[0002] With the rapid development of the port and shipping industry, the rapid updating and timely acquisition of professional knowledge has become crucial. Intelligent question-and-answer systems are increasingly being used in the industry, improving the efficiency and accuracy of information retrieval through automation and intelligence, thereby enhancing the efficiency of staff in obtaining port and shipping data.

[0003] In existing technologies, data in the port and shipping industry involves a large number of professional terms and complex operational procedures. Traditional knowledge management systems often rely on manual maintenance, resulting in untimely knowledge updates, low review efficiency, and difficulty in ensuring content accuracy. Furthermore, most existing intelligent question-answering systems are rule-based or simple database query systems. Rule-based question-answering systems use predefined rules and templates for question-and-answering, but their rule updates are complex and inflexible, making it difficult to handle complex and ever-changing professional knowledge, and they have poor understanding of user input. Database query systems rely on fixed databases for queries, and they cannot intelligently review and add new data during use, thus making it difficult to guarantee content accuracy and timeliness.

[0004] Currently, no effective solutions have been proposed to address the problems of existing question-and-answer systems' poor ability to understand user-input text and their inability to intelligently review and approve data before it is entered into the database. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method for constructing a question-answering system based on a large language model and a question-answering system, so as to at least solve the problems of existing question-answering systems having poor ability to understand user input text and the inability to intelligently review and update database data.

[0006] The embodiments of the present invention provide the following technical solutions:

[0007] This invention provides a method for constructing a question-answering system based on a large language model, comprising:

[0008] Acquire port and shipping data, port and shipping industry knowledge and industry standards, and generate data for storage after preprocessing the port and shipping data;

[0009] Based on the port and shipping industry knowledge and standards, a generative large language model is used to generate review rules. The review rules work by combining retrieval enhancement generation with prompt words.

[0010] Based on the review rules, the large language model is used to generate a review module. The review module includes review knowledge text and preset review prompt words. The review module can perform retrieval enhancement generation on the data in the database based at least on the port and shipping industry knowledge, the port and shipping industry standards and the review knowledge text to obtain review retrieval information corresponding to the data in the database.

[0011] Based on the audit module and the audit retrieval information, multiple audit language models are used to audit the data entering the database. When all of the multiple audit language models pass the audit, the data entering the database is entered into the port and shipping database.

[0012] The generated large language model is used to build a question-answering engine, and the user input is understood based on the question-answering engine and the port and shipping database, and the corresponding user output is output.

[0013] Furthermore, based on the review module and the review retrieval information, the multiple review language models are used to review the data entering the database. If at least one of the review language models fails the review, the data entering the database is transferred to manual review.

[0014] Furthermore, the question-answering system construction method also includes:

[0015] Based on the review rules, the generated large language model is used to generate an error identification and completion module. The error identification and completion module includes error identification and completion text and preset completion prompts. The error identification and completion module can perform retrieval enhancement generation on the input data based at least on the port and shipping industry knowledge, the port and shipping industry standards and the error identification and completion text to obtain completion retrieval information corresponding to the input data.

[0016] Based on the error identification and completion module and the completion retrieval information, the completion language model is used to identify and correct the error information in the data entering the database and / or complete the correct information in the data entering the database, and the data entering the database after identification, correction and / or completion is entered into the port and shipping database.

[0017] Furthermore, a few sample prompts are used to add prompt examples to the preset audit prompts and / or preset completion prompts.

[0018] Furthermore, the step of using the generated large language model to build a question-answering engine, and understanding user input based on the question-answering engine and the port and shipping database, and outputting the corresponding user output includes:

[0019] The question-answering engine is constructed using the generative large language model, and the question-answering engine includes question-answering knowledge text and question-answering prompt words;

[0020] Based on the question-answering engine, the question-answering large language model is used to obtain and understand user input. The question-answering large language model generates a connection to the port and shipping database through retrieval enhancement to obtain input retrieval information corresponding to the user input.

[0021] Based on the user input and the input retrieval information corresponding to the user input, the question-answering big language model outputs the user output corresponding to the user input.

[0022] Furthermore, after constructing a question-answering engine using the aforementioned generative large language model, understanding user input based on the question-answering engine and the port and shipping database, and outputting the corresponding user output, the process further includes:

[0023] Obtain negative user feedback and the corresponding interaction data;

[0024] The negative user feedback, the interaction data, and the corresponding data entered into the database are manually reviewed, and the corresponding data is re-entered into the port and shipping database after the review is approved.

[0025] Furthermore, after constructing a question-answering engine using the aforementioned generative large language model, understanding user input based on the question-answering engine and the port and shipping database, and outputting the corresponding user output, the process further includes:

[0026] Detect and identify updated data, which is data that needs to be updated into the port and shipping database;

[0027] Based on the review module and the review retrieval information, the updated data is reviewed using multiple review language models. If all multiple review language models pass the review, the updated data is entered into the port and shipping database; or

[0028] Based on the error identification and completion module and the completion retrieval information, the completion language model is used to identify and correct errors in the updated data and / or complete correct information in the updated data, and the identified, corrected and / or completed updated data is entered into the port and shipping database.

[0029] Furthermore, after entering the updated data into the port and shipping database, the process also includes:

[0030] Record the time and content of the updated data being entered into the port and shipping database.

[0031] Furthermore, the step of acquiring port and shipping data and generating data for storage after preprocessing the port and shipping data includes:

[0032] The port and shipping data is automatically collected through multiple channels, including one or more of the following: port and shipping professional literature, databases, websites, and internal company documents.

[0033] The port and shipping data is deduplicated, completed, and formatted to ensure its integrity and consistency.

[0034] Based on the domain and theme of the port and shipping data, the port and shipping data is classified and organized to generate the data to be stored in the database.

[0035] This invention also provides a question-answering system based on a large language model, comprising:

[0036] A knowledge base construction module is used to acquire port and shipping data, port and shipping industry knowledge and industry standards, and to preprocess the port and shipping data to generate data to be stored in the database.

[0037] The knowledge base review module, based on the port and shipping industry knowledge and standards, uses a generative large language model to generate review rules. These review rules operate using a combination of retrieval enhancement generation and prompt words. Based on these review rules, the generative large language model generates a review module. This module includes review knowledge text and preset review prompt words. The review module can perform retrieval enhancement generation on the input data based at least on the port and shipping industry knowledge, standards, and the review knowledge text to obtain review retrieval information corresponding to the input data. Based on the review module and the review retrieval information, multiple review large language models are used to review the input data. When all multiple review large language models pass the review, the input data is entered into the port and shipping database.

[0038] The question-answering engine building module is used to build a question-answering engine using the generated large language model, and to understand user input based on the question-answering engine and the port and shipping database, and output the corresponding user output.

[0039] Compared with the prior art, the beneficial effects that the at least one technical solution adopted in the embodiments of the present invention can achieve include at least:

[0040] This invention discloses a question-answering system construction method based on a large language model. It generates review rules based on port and shipping industry knowledge and standards, and then uses the generated large language model to create a review module based on these rules. This review module can retrieve and enhance the input data based on port and shipping industry knowledge and standards to obtain review retrieval information. When the input data and review retrieval information are input into the review large language model, the model can more accurately understand the input data and perform precise review under the guidance of preset review prompts. Finally, after multiple review large language models have approved the data, it is entered into the port and shipping database. This achieves intelligent review of the input data, solving the problem that existing question-answering systems cannot achieve intelligent review and update of database data, and improving the accuracy of input data review.

[0041] Furthermore, by constructing a question-answering engine using a large generative language model and question-answering prompts, and by using retrieval enhancement to generate a related port and shipping database, the question-answering engine's ability to understand user input can be improved. Moreover, since the question-answering engine generates a related external port and shipping database based on retrieval enhancement, the correctness of the output answer is also improved. This solves the problem in the prior art that question-answering based on predefined rules and templates is difficult to handle complex and ever-changing professional knowledge and has poor understanding of user input text. Attached Figure Description

[0042] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 This invention provides a method for constructing a question-answering system based on a large language model.

[0044] Figure 2 This is a specific process for constructing a question-answering system based on a large language model, according to an embodiment of the present invention.

[0045] Figure 3 This is a structural block diagram of a question-answering system based on a large language model according to an embodiment of the present invention. Detailed Implementation

[0046] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0047] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0048] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this application, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number and aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0049] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. The drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0050] Additionally, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that practice can be carried out without these specific details.

[0051] With the rapid development of the port and shipping industry, the rapid updating and timely acquisition of professional knowledge has become particularly important. Intelligent question-and-answer systems are gradually being applied in the port and shipping industry to improve the efficiency and accuracy of information acquisition through automation and intelligent means.

[0052] However, existing technologies still have many problems in the construction, review and updating of knowledge bases, as well as the level of intelligence of question-answering systems. For example, knowledge bases cannot be updated automatically, requiring manual maintenance, review and updating of data entering the knowledge base. Question-answering systems have poor understanding of user questions, making it difficult to fully understand user questions, or they cannot flexibly match relevant answers to users after understanding user questions. Therefore, the existing question-answering systems are not very effective and usually require manual solutions.

[0053] Because the port and shipping industry involves a large number of technical terms and complex operational procedures, traditional knowledge management systems often rely on manual maintenance. However, manual review of data leads to untimely knowledge updates, low review efficiency, and difficulty in ensuring content accuracy (if the review is not conducted by experts within the port and shipping industry, issues such as content matching errors or incorrect updates may occur). Given the complexity, rapid pace of knowledge updates, and high difficulty of understanding in the port and shipping industry, an intelligent question-and-answer system is required to possess efficient knowledge updating, comprehension, and review capabilities to meet the industry's need for timely and accurate access to professional knowledge.

[0054] Because knowledge in the port and shipping sector is not widely known to the public, and is complex, specialized, and diverse, and because most current question-and-answer systems in the port and shipping sector are built on predefined rules and limited databases, their ability to understand user questions is poor. Furthermore, due to the limited data in their databases, existing question-and-answer systems are also unable to provide users with correct answers.

[0055] Specifically, existing question-answering systems have the following problems:

[0056] (1) High cost of manual review: The review and update of existing database content in the port and shipping sector requires a large amount of manpower, which is costly and inefficient;

[0057] (2) Low accuracy of content: Due to the subjectivity and tediousness of manual review, errors or omissions are likely to occur in the data entering the database, which will result in the subsequent question and answer system being unable to answer user questions correctly;

[0058] (3) Lagging updates: Database updates rely on manual operation, making it difficult to reflect the latest industry trends and knowledge in a timely manner;

[0059] (4) Lack of automatic replenishment: The existing database is difficult to automatically replenish and update content, resulting in an incomplete knowledge base.

[0060] Therefore, the aforementioned problems have limited the effectiveness of intelligent question-and-answer systems in the port and shipping industry to some extent, making it difficult to meet users' needs for efficient and accurate information acquisition.

[0061] The following are examples of existing intelligent question-answering systems. These systems generally include rule-based question-answering systems and simple database query systems. Rule-based question-answering systems typically use predefined rules and templates for question-and-answering, but their rule updates are complex and inflexible, making it difficult to handle complex and ever-changing professional knowledge. Database query systems, on the other hand, rely on fixed databases for queries, failing to achieve intelligent knowledge updates and review, and making it difficult to guarantee the accuracy and timeliness of the content.

[0062] Based on this, the embodiments of this specification propose a processing solution: such as Figure 1 As shown, this invention provides a method for constructing a question-and-answer system based on a large language model. It automatically acquires port and shipping data, generates database data after preprocessing, and then generates review rules applicable to the port and shipping field based on port and shipping industry knowledge and standards. This enables the subsequent review module to retrieve and enhance the database data to obtain review retrieval information. When the review module is used to review the large language model, the large language model can obtain the database data and corresponding review retrieval information to accurately understand the data. Finally, guided by preset review prompts, the large language model accurately reviews the database data and, after approval, enters it into the database. This not only improves the accuracy of database data review but also solves the problem that existing question-and-answer systems cannot achieve intelligent review and update of database data.

[0063] Furthermore, this invention constructs a question-answering engine based on a generative large language model and generates a related port and shipping database based on retrieval enhancement. This enables the question-answering large language model to accurately understand user input and accurately output the corresponding output based on retrieval enhancement, thereby improving the accuracy of user question-answering and reducing the cost of manual review.

[0064] The technical solutions provided by the various embodiments of this application are described below with reference to the accompanying drawings.

[0065] Example 1

[0066] This invention provides a method for constructing a question-answering system based on a large language model, such as... Figure 1 As shown, the method includes:

[0067] Step S102: Obtain port and shipping data, port and shipping industry knowledge and port and shipping industry standards, and generate data for storage after preprocessing the port and shipping data.

[0068] When acquiring port and shipping data, data acquisition units can be used to automatically collect port and shipping data, such as using Python scraping tools to automatically scrape port and shipping data.

[0069] Among them, data scraping tools can be used to automatically acquire port and shipping industry knowledge and standards.

[0070] Specifically, step S102 includes:

[0071] Step S102a: Automatically collect port and shipping data through multiple channels, including one or more of the following: port and shipping professional literature, databases, websites, and internal company documents;

[0072] Step S102b: Deduplicatize, complete, and format the port and shipping data to ensure its integrity and consistency.

[0073] Step S102c: Classify and organize the port and shipping data according to the field and theme to generate data for storage.

[0074] Among them, port and shipping data can be automatically collected through multiple channels via the data acquisition unit.

[0075] Among them, internal company documents are knowledge documents published by company employees. These documents can be reviewed by experts or not.

[0076] By preprocessing the collected data through steps S102b and S102c, port and shipping data with relatively high completeness and consistency can be obtained. Finally, the port and shipping data is classified and organized according to the field and theme to generate data for storage, which facilitates the subsequent review of the data.

[0077] Step S102 enables the automatic collection and crawling of port and shipping data in the port and shipping sector. After a series of processing steps, the data is formatted to be entered into the port and shipping database, which also facilitates intelligent review and entry of the data into the database.

[0078] Step S104: Based on port and shipping industry knowledge and standards, generate audit rules using a large language model. The audit rules are generated by combining retrieval enhancement with prompt words.

[0079] The review rules are derived from industry knowledge characteristics (such as international ocean organization standards, TOS system information of a certain terminal, technical specifications, operating procedures, etc.) and port and shipping industry standards, which are manually reviewed by experts to ensure the professionalism and accuracy of the rules. This ensures that when the enhanced search function searches the data in the database, the reviewed search information retrieved is strongly related to the port and shipping industry and conforms to the standards and characteristics of the port and shipping industry.

[0080] Since the amount of existing professional knowledge data in the port and shipping field is insufficient to achieve the scale of model fine-tuning, the review rules are generated by combining search enhancement with review prompts. This ensures that the subsequent review module can still perform the review task well and maintain a certain level of accuracy and relevance, even with limited data in the port and shipping field.

[0081] Specifically, retrieval enhancement and generation mainly include three processes: retrieval, enhancement, and generation. Retrieval involves obtaining relevant information from an external knowledge base based on the user's query. Specifically, the user's query is converted into a vector using an embedding model for comparison with relevant knowledge stored in a vector database. Similarity search is used to find the top K most relevant data points. Enhancement involves embedding the user's query and the retrieved relevant knowledge into a pre-defined prompt template. Generation involves inputting the enhanced prompt content into a large language model to generate the desired output. Therefore... Search enhancement generation, after obtaining user input, can perform searches within its associated database and generate the results. The retrieved content is then input into the large language model. This allows large language models to more accurately understand user input based on the information they acquire.

[0082] The prompt words are used to guide the generation of review rules by the large language model, so as to help the large language model clarify the direction of the generated content based on the knowledge and standards of the port and shipping industry, and ensure that the generated review rules can effectively cover the industry standards and knowledge of the port and shipping industry, so as to adjust and optimize the generation process.

[0083] Step S104 works by combining search enhancement generation with prompt words. Even when the amount of data in the port and shipping industry is insufficient to reach the level of model fine-tuning, it can still generate review rules well and maintain a certain level of accuracy and relevance. This overcomes the problem of insufficient data that may be caused by relying solely on the model's inherent knowledge, further improving the quality and practicality of the review rule generation. Moreover, by using search enhancement generation to search within the knowledge and standards of the port and shipping industry, the retrieved information can be more in line with the port and shipping industry.

[0084] Step S106: Based on the review rules, use the large language model to generate the review module. The review module includes review knowledge text and preset review prompt words. The review module can perform retrieval enhancement on the data in the database based on at least port and shipping industry knowledge, port and shipping industry standards and review knowledge text to obtain review retrieval information corresponding to the data in the database.

[0085] The audit module is used to audit the data entering the database when it is applied to the large language model for auditing.

[0086] Among them, the audit knowledge text is a knowledge document that has been manually reviewed by experts and is related to the port and shipping industry. It can also be a knowledge document of a specific sub-sector within the port and shipping industry. The audit knowledge text can also be used to provide some data sources for the audit language model when the audit module is applied to the audit language model, or to provide the audit module with the context information required during the actual audit.

[0087] The preset audit prompts are used to guide the large language model to audit the data when the audit module is applied to the large language model. The preset audit prompts are manually set and can be flexibly set according to needs, as long as their purpose is to guide the large language model to audit the data.

[0088] The following is an example of an approval module generated by the large language model based on preset approval prompts and approval rules:

[0089]

[0090] <context>

[0091] [context]

[0092] < / context>

[0093] You are now an expert assistant in the port and shipping industry. You need to complete the review task using the given context information.

[0094] Your task is to verify the accuracy of user input.

[0095] If you are unsure whether your input is correct, simply reply "I'm not sure".

[0096] If you are sure that the content you entered is correct, you only need to reply "Correct".

[0097] If you are certain that there is an error in your input, you only need to reply "Error exists".

[0098] "

[0099] In the example of climbing a tree, the text portion is a preset review prompt, and the context information is the review knowledge text, which is a knowledge document reviewed by experts.

[0100] Specifically, after the review module obtains the data, since the review rules are based on port and shipping industry knowledge and standards and work by combining search enhancement generation with prompt words, the review module can enhance the search of the data within the port and shipping industry knowledge and standards through search enhancement generation, and generate review search information. Then, the review module sends the data and the review search information to the corresponding review language model, so that the review language model can accurately understand the data in the port and shipping field, and can generate output corresponding to the data or generate explanations corresponding to the data based on the port and shipping industry knowledge and standards associated with the search enhancement generation, or accurately review the data under the effect of preset review prompt words.

[0101] The retrieval enhancement generation includes input enhancement, retrieval enhancement, generator enhancement, result enhancement, and RAG process enhancement. Therefore, after the retrieval enhancement generation works in conjunction with prompt words and is associated with databases such as port and shipping industry knowledge, port and shipping industry standards, and audit knowledge texts, the audit module can obtain audit retrieval information that is related to the data in the database and conforms to the characteristics of port and shipping industry knowledge and standards based on the retrieval enhancement generation. This provides more accurate and reliable port and shipping domain knowledge for the audit big language model, reduces the possibility of generating illusions, and can also maintain the timeliness and accuracy of the output by accessing the latest external knowledge base.

[0102] Step S106 uses the review rules to generate a review module using a large language model. The review module can perform retrieval enhancement on the input data to obtain review retrieval information, and send the retrieval enhancement-generated input data and the review retrieval information together to the review large language model, so that the review large language model can understand the input data more accurately and avoid the problem of the review large language model having illusions after obtaining the input data.

[0103] Step S108: Based on the audit module and audit retrieval information, multiple audit language models are used to audit the data entering the database. When all multiple audit language models pass the audit, the data entering the database is entered into the port and shipping database.

[0104] When the review module is integrated into the review language model, the review module sends the input data and the review retrieval information corresponding to the input data to the review language model. The review language model can understand the input data based on the review retrieval information, provided that it complies with the knowledge and standards of the port and shipping industry. Under the guidance of the preset review prompts in the review module, it can review the input data by combining the review knowledge text, the review retrieval information, and its own retrieval data.

[0105] After the large language model receives the input data and retrieval information sent by the review module, it can more accurately understand the input data. As a result, the large language model can perform accurate queries within its associated database and then review the input data under the guidance of the preset review prompts in the review module.

[0106] Among them, multiple review language models can be different from each other, and a language model can be the same as any review language model.

[0107] For example, when there are multiple large language models for review, including models such as Qwen, LLama, gemma, and Yi, the generated large language model can be any one of the modules such as Qwen, LLama, gemma, and Yi.

[0108] This approach integrates the review module into multiple large review language models, allowing for simultaneous review of incoming data. Once all large review language models have approved the data, it is then entered into the database. This effectively addresses the illusion problem of large models and improves the accuracy of the review process.

[0109] Therefore, in step S108, the review module uses a voting mechanism with multiple review language models, and only data that passes unanimously can be directly entered into the port and shipping database.

[0110] Specifically, the system uses multiple large-scale review language models to simultaneously review the incoming data. Once all the large-scale review language models have approved the data, the data is directly entered into the database, thus achieving intelligent data entry. Because both the review module and the large-scale review language models have good understanding and logical abilities, and because they are linked to port and shipping industry knowledge and standards through enhanced retrieval, the system can ensure the accuracy of the incoming data. This solves the problem of manual review of incoming data in existing technologies and improves the efficiency of data entry.

[0111] Furthermore, based on the review module and review retrieval information, multiple review big language models are used to review the data entering the database. If at least one review big language model fails the review, the data entering the database is transferred to manual review. After the manual review is passed, the data entering the database is added to the port and shipping database.

[0112] If the manual review fails, the data to be entered into the database will be discarded.

[0113] By submitting data that failed to pass the review of multiple large language models to human reviewers or relevant experts for verification, the accuracy of knowledge can be ensured, the illusion problem of large language models can be avoided, and the correctness of the answers output by the subsequent question-answering engine can be improved.

[0114] In step S108, multiple auditing language models are used to audit the incoming data simultaneously, and manual auditing is combined with the auditing of the data in the multiple auditing language models that show discrepancies. This ensures the correctness of the data entering the port and shipping database. Furthermore, the automatic auditing of the incoming data using multiple auditing language models also enables intelligent updating and auditing of the data in the port and shipping database.

[0115] In some embodiments, the question-answering system construction method further includes:

[0116] Step S110: Based on the review rules, use the large language model to generate an error identification and completion module. The error identification and completion module includes error identification and completion text and preset completion prompts. The error identification and completion module can at least perform retrieval enhancement on the data in the database based on port and shipping industry knowledge, port and shipping industry standards and error identification and completion text to obtain completion retrieval information corresponding to the data in the database.

[0117] Step S112: Based on the error identification and completion module and the completion retrieval information, use the completion large language model to identify and correct erroneous information in the data and / or complete accurate information in the data, and then enter the identified, corrected and / or completed data into the port and shipping database.

[0118] The error identification and completion module is used to enhance the retrieval of the input data based on port and shipping industry knowledge, port and shipping industry standards, and error identification and completion text to obtain completion retrieval information. Then, the input data and the corresponding completion retrieval information are sent to the completion language model. The completion language model accurately understands the input data based on the error identification and completion text and completion retrieval information, and performs error identification or completion of the input data under the guidance of preset completion prompt words.

[0119] The following is an error recognition and completion module generated by the large language model based on preset completion prompts and combined with review rules:

[0120]

[0121] <context>

[0122] [context]

[0123] < / context>

[0124] You are now an expert assistant in the port and shipping industry. You need to complete tasks using the given context information.

[0125] First, you need to identify and correct errors in the user's input.

[0126] Next, you need to complete the exact information in the input content.

[0127] If the user modifies the input, you need to output the modified content;

[0128] If the user input remains unchanged, you only need to output the user's input.

[0129] And answer in Chinese.

[0130] "

[0131] The text portion contains preset completion prompts, while the context information contains error-identified completion text.

[0132] In step S112, completing the large language model may be the same as or different from generating or reviewing the large language model.

[0133] The error identification and completion module is generated based on audit rules, which are based on search enhancement and prompt words. As a result, the completed search information obtained by the error identification and completion module after enhancing the search of the input data is highly correlated with the input data and conforms to port and shipping industry knowledge and standards. Therefore, after the completion search information and input data are obtained by the completion language model, the module can identify and correct errors in the input data or complete the correct information in the input data, and then enter the error-identified or completed input data into the port and shipping database.

[0134] In some of these embodiments, few-sample prompts can be used to add prompt examples to preset review prompts and / or preset completion prompts, so that the completion or review of the large language model can be significantly improved in its awareness of its own role positioning, thereby effectively reducing the illusion problem of the large language model and improving the accuracy of review, error identification and correction, completion and question answering.

[0135] Steps S110-S112 enable intelligent error identification and completion of the data entering the database, further improving the accuracy of the data entered into the port and shipping database.

[0136] Step S114: Use a generative large language model to build a question-answering engine, and understand user input based on the question-answering engine and the port and shipping database, and output the corresponding user output.

[0137] After building the port and shipping database, a question-answering engine can be built using a large generative language model to engage in question-and-answer sessions with users.

[0138] Specifically, after the question-answering engine obtains and understands the user input, it can match the corresponding user output to the user input based on the port and shipping database. Since the question-answering engine is built on a large language model, it has good text understanding and generation capabilities, thus providing users with user output with a high accuracy rate.

[0139] In some embodiments, when the question-answering engine cannot fully understand the user's input, it can continue to ask the user questions so that the user can continue to enter user questions until the question-answering engine fully understands the user's questions.

[0140] Specifically, step S114 includes:

[0141] Step S114a: Use a generative large language model to build a question-answering engine, which includes question-answering knowledge text and question-answering prompt words.

[0142] The question-and-answer knowledge text includes at least all the data entered into the port and shipping database; the question-and-answer prompt words are used to guide the question-and-answer language model to answer user input when the question-and-answer engine is used in the question-and-answer language model.

[0143] The following is an example of a question-answering engine generated from a large language model:

[0144]

[0145] <context>

[0146] [context]

[0147] < / context>

[0148] You are now an expert assistant in the port and shipping industry. You need to respond to user input using the given context information.

[0149] If you don't know the answer, just say you don't know.

[0150] If you are unsure of the answer, ask the user for more detailed information.

[0151] And answer the user's question based on the language used.

[0152] "

[0153] The text portion consists of question-and-answer prompts, while the context information is the question-and-answer knowledge text, which includes at least all knowledge documents in the knowledge base.

[0154] Step S114b: Based on the question-answering engine, the question-answering big language model is used to obtain and understand user input. The question-answering big language model generates a related port and shipping database through retrieval enhancement to obtain input retrieval information corresponding to the user input.

[0155] Among them, the question-answering big language model generates a related port and shipping database through retrieval enhancement. Its retrieval enhancement generation can retrieve the input retrieval information based on the port and shipping database. Then, the retrieval enhancement generation sends the input retrieval information to the question-answering big language model so that the question-answering big language model can accurately understand the user input by combining the input retrieval information and the user input.

[0156] Step S114c: Based on the user input and the input retrieval information corresponding to the user input, the question-answering big language model outputs the user output corresponding to the user input.

[0157] After the question-answering big language model obtains user input and input retrieval information and accurately understands the user input, it can obtain user output based on question-answering knowledge text and its own retrieval, and output the user output under the influence of question-answering prompt words.

[0158] Steps S114a to S114c first use retrieval enhancement generation to enhance the user input to obtain input retrieval information, thereby enabling the question-answering big language model to accurately understand the user input based on the input retrieval information, thus improving the question-answering big language model's ability to understand user input and also improving the answer accuracy of the question-answering engine.

[0159] In summary, the question-answering engine built on the large language model in step S110 supports natural language questioning and answering. With the powerful semantic understanding capability of the large language model, it can accurately understand the user's input. At the same time, by using retrieval enhancement generation technology to associate with the private port and shipping database, it can also provide users with professional and accurate answers.

[0160] Furthermore, the methods for building question-answering systems also include:

[0161] Step S116: Obtain user negative feedback and the corresponding interaction data;

[0162] Step S118: Manually review user negative feedback, interaction data, and corresponding data entering the database, and re-enter the corresponding data into the port and shipping database after the review is approved.

[0163] Among them, negative user feedback refers to negative evaluations of the answers output by the question-and-answer engine by users, and interaction data refers to the question-and-answer data of the question-and-answer engine, which includes user question data and question-and-answer engine answer data.

[0164] Users can provide positive or negative feedback on the answers output by the question-and-answer engine. For questions that users frequently report, it indicates that the question-and-answer engine is experiencing high-frequency triggering of illusions. After being reviewed and confirmed by human experts, these questions can be added to the database for correction to improve the accuracy of subsequent questions and answers.

[0165] Furthermore, the question-answering system construction method also includes:

[0166] Step S202: Detect and identify the updated data, which is the data that needs to be updated in the port entry database;

[0167] Step S204: Based on the review module and review retrieval information, multiple review language models are used to review the updated data. If all multiple review language models pass the review, the updated data is entered into the port and shipping database; or

[0168] Step S206: Based on the error identification and completion module and the completion retrieval information, use the completion large language model to identify and correct error information in the updated data generated after retrieval enhancement and / or complete the correct information in the updated data, and enter the identified, corrected and / or completed updated data into the port and shipping database.

[0169] After building the port and shipping database and question-and-answer engine, it is also necessary to detect the latest developments, information and internal knowledge documents in the port and shipping field in order to identify the knowledge content that needs to be updated.

[0170] In step S202, the data acquisition unit can be used to detect and identify updated data in the port and shipping sector, such as the port and shipping database, that needs to be updated.

[0171] Steps S204 and S206 are used to review, identify, or complete the updated data to ensure its accuracy.

[0172] Furthermore, after entering the updated data into the port and shipping database, the process also includes:

[0173] Step S208: Record the time and content of the updated data being entered into the port and shipping database.

[0174] Step S208 records the time and content of the updated data entered into the port and shipping database, thereby facilitating the management of different versions of the port and shipping database and recording the content and time of each update for traceability and management.

[0175] like Figure 3As shown, this embodiment of the invention also provides a question-answering system based on a large language model, including a knowledge base construction module 10, a knowledge base review module 20, and a question-answering engine construction module 30. The knowledge base construction module 10 is used to acquire port and shipping data, port and shipping industry knowledge, and industry standards, and to generate data for storage after preprocessing the port and shipping data. The knowledge base review module 20 is used to generate review rules based on port and shipping industry knowledge and standards using a generative large language model. The review rules work by combining retrieval enhancement generation with prompt words. Based on the review rules, a review module is generated using a generative large language model. The review module includes review knowledge text and preset review prompt words. The review module can perform retrieval enhancement generation on the data for storage based on at least port and shipping industry knowledge, port and shipping industry standards, and review knowledge text to obtain review retrieval information corresponding to the data for storage. Based on the review module and the review retrieval information, multiple review large language models are used to review the data for storage. When multiple review large language models pass the review, the data for storage is entered into the port and shipping database. The question answering engine construction module 30 is used to build a question answering engine using a generative large language model, and to understand user input based on the question answering engine and the port and shipping database, and output the corresponding user output.

[0176] The knowledge base construction module 10 further includes a data acquisition unit, a data cleaning unit, and a knowledge classification unit. The data acquisition unit automatically collects port and shipping professional knowledge data from relevant literature, websites, and the company's internal knowledge platform. The data cleaning unit performs deduplication and formatting on the collected data to ensure its usability and consistency. The knowledge classification unit categorizes and organizes the data according to the domain and theme of port and shipping professional knowledge for subsequent querying and management.

[0177] The knowledge review module is used to intelligently review and update the content in the knowledge base to ensure the accuracy and timeliness of the knowledge. It includes a review rule unit and an automatic review unit. The review rule generation unit generates review rules based on the port and shipping industry knowledge and standards using a generative large language model. These review rules operate using a combination of retrieval enhancement generation and prompt words. The automatic review unit generates a review module based on the review rules and the generative large language model. The review module includes review knowledge text and preset review prompt words. The review module can perform retrieval enhancement generation on the input data based at least on the port and shipping industry knowledge, the port and shipping industry standards, and the review knowledge text to obtain review retrieval information corresponding to the input data. Based on the review module and the review retrieval information, multiple review large language models are used to review the input data. When all multiple review large language models pass the review, the input data is entered into the port and shipping database.

[0178] The automatic review unit is also used to automatically review the incoming data using a large language model, identify and correct errors, update outdated content, and supplement missing contextual information in the incoming data.

[0179] The automatic review unit is also used to determine whether a manual review is needed by using multiple language models and a final voting mechanism, thus resolving the illusion problem caused by a single language model.

[0180] The question-answering system building module is used to establish an intelligent question-answering engine, enabling efficient interaction with users. This module includes a question-answering engine unit, a knowledge matching unit, and a learning optimization unit. The question-answering engine unit is used to build a question-answering engine based on LLM, supporting natural language questioning and answering. The knowledge matching unit matches user questions with content in the knowledge base, providing the most relevant and accurate answers. The learning optimization unit continuously optimizes the performance and accuracy of the question-answering engine through user feedback and interaction data.

[0181] The question-answering system construction device also includes a knowledge update module, which is used to realize the timed updating and automatic replenishment of knowledge base content. This module includes a timed monitoring unit, a content generation unit, and a version control unit. Specifically, the timed monitoring unit monitors the latest developments and information in the port and shipping sector and identifies knowledge content that needs updating; the content generation unit uses a large language model to generate new knowledge entries and automatically adds them to the knowledge base; and the version control unit manages different versions of the knowledge base, recording the content and time of each update for traceability and management.

[0182] Through the collaborative work of the above-mentioned units, this embodiment of the invention constructs an efficient, accurate, and intelligent port and shipping professional knowledge question-and-answer system, which greatly improves the human-computer interaction experience, solves many problems in the prior art, and meets the port and shipping industry's needs for efficient acquisition and updating of professional knowledge.

[0183] The question-answering system construction method and apparatus based on a large language model according to embodiments of the present invention have the following advantages compared with the prior art:

[0184] 1. Reduce manual review costs

[0185] By introducing an intelligent review module, this invention achieves automatic review and updating of knowledge base content, significantly reducing the workload and cost of manual review. Utilizing the natural language processing capabilities of a large language model, errors in the knowledge base can be efficiently identified and corrected, further reducing reliance on manual review.

[0186] 2. Improve review efficiency

[0187] The automatic review module of this invention can quickly and efficiently review knowledge base content, ensuring the timeliness and accuracy of the knowledge. Compared with the traditional method of relying on manual review item by item, the review module can complete the review of a large amount of data in a short time, improving the overall review efficiency.

[0188] 3. Enhance content accuracy

[0189] By combining automated review and manual verification, the knowledge review module of this invention ensures high accuracy of the knowledge base content. The automated review module first uses multiple models for preliminary review, identifies complex or difficult issues based on a voting mechanism, and then the manual verification unit processes the data, ensuring the correctness and reliability of the content.

[0190] 4. Enable automatic replenishment and updates.

[0191] The knowledge update module of this invention can automatically monitor the latest developments in the port and shipping industry and generate new knowledge entries using LLM, achieving real-time updates and automatic replenishment of the knowledge base content. This solves the problem of automatic replenishment during knowledge updates in existing technologies, ensuring that the knowledge base always reflects the latest professional information.

[0192] 5. Improve user interaction experience

[0193] By constructing an intelligent question-answering system (question-answering engine), this invention significantly enhances the user's interactive experience. The question-answering engine, based on a large language model, can understand the natural language questions posed by users and provide accurate and professional answers. A knowledge matching unit ensures the relevance and accuracy of the answers, while a learning optimization unit continuously improves system performance through user feedback, further enhancing the user experience.

[0194] 6. Reduce maintenance costs

[0195] By automating knowledge review and updates, this invention reduces the frequency and cost of manual maintenance. The system can automatically monitor and update knowledge base content, reducing manual intervention and improving system maintenance efficiency and effectiveness.

[0196] In summary, this invention provides an intelligent, automated, efficient, and accurate port and shipping professional knowledge question-and-answer system, which significantly improves the performance of existing technologies and solves problems such as high manual review costs, low content accuracy, lagging knowledge updates, and lack of automatic supplementation. It has broad application prospects and significant technical advantages in intelligent question-and-answer applications in the port and shipping industry.

[0197] Example 2

[0198] This embodiment is a specific implementation of Embodiment 1, such as... Figure 2 As shown, the specific steps are as follows:

[0199] Step S401: Gather professional knowledge;

[0200] Step S402: Review / supplement the knowledge content of the multi-agent (large language model);

[0201] Step S403: Manually review the content that fails the initial review;

[0202] Step S404: Construction of the professional knowledge base is complete;

[0203] Step S405: Construct an intelligent question-answering system based on RAG technology;

[0204] Step S406: Construct domain experts for review prompts;

[0205] Step S407: System Deployment.

[0206] In this specification, the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the product embodiments described later, since they correspond to the methods, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions in the system embodiments.

[0207] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for constructing a question-answering system based on a large language model, characterized in that, include: Acquire port and shipping data, port and shipping industry knowledge, and port and shipping industry standards, and generate data for storage after preprocessing the port and shipping data; Based on the aforementioned port and shipping industry knowledge and standards, a generative language model is used to generate review rules. These review rules operate by combining retrieval-enhanced generation with prompt words. The prompt words are used to guide the generative language model in generating the review rules, helping the model to clarify the direction of the generated content based on port and shipping industry knowledge and standards, thus ensuring that the generated review rules can effectively cover industry standards and knowledge in the port and shipping field. Based on the aforementioned review rules, a review module is generated using the generative large language model. The review module includes review knowledge text and preset review prompts. The review module can perform retrieval enhancement on the input data based at least on the port and shipping industry knowledge, the port and shipping industry standards, and the review knowledge text to obtain review retrieval information corresponding to the input data. The review retrieval information conforms to the port and shipping industry knowledge and the port and shipping industry standards. The preset review prompts are used to guide the review large language model to review the input data when the review module is applied to the review large language model. Based on the review module and the review retrieval information, multiple different review language models are used to review the data entering the database. When all of the multiple review language models pass the review, the data entering the database is entered into the port and shipping database. The review module sends the data entering the database and the corresponding review retrieval information to the review language model. The review language model understands the data entering the database based on the review retrieval information and, guided by the preset review prompt words, reviews the data entering the database by combining the review knowledge text, the review retrieval information, and its own retrieval data. The generated large language model is used to build a question-answering engine, and the user input is understood based on the question-answering engine and the port and shipping database, and the corresponding user output is output.

2. The question-answering system construction method according to claim 1, characterized in that, Based on the review module and the review retrieval information, the data entering the database is reviewed using the multiple review language models. If at least one of the review language models fails the review, the data entering the database is transferred to manual review.

3. The question-answering system construction method according to any one of claims 1 or 2, characterized in that, Also includes: Based on the review rules, the generated large language model is used to generate an error identification and completion module. The error identification and completion module includes error identification and completion text and preset completion prompts. The error identification and completion module can perform retrieval enhancement generation on the input data based at least on the port and shipping industry knowledge, the port and shipping industry standards and the error identification and completion text to obtain completion retrieval information corresponding to the input data. Based on the error identification and completion module and the completion retrieval information, the completion large language model is used to identify and correct the error information in the data entering the database and / or complete the correct information in the data entering the database, and the data entering the database after identification, correction and / or completion is entered into the port and shipping database.

4. The question-answering system construction method according to claim 3, characterized in that, Use a few samples to add example prompts to the preset audit prompts and / or preset completion prompts.

5. The question-answering system construction method according to claim 3, characterized in that, The step of using the generated large language model to build a question-answering engine, and understanding user input based on the question-answering engine and the port and shipping database, and outputting the corresponding user output includes: The question-answering engine is constructed using the generative large language model, and the question-answering engine includes question-answering knowledge text and question-answering prompt words; Based on the question-answering engine, a question-answering large language model is used to obtain and understand user input. The question-answering large language model generates a connection to the port and shipping database through retrieval enhancement to obtain input retrieval information corresponding to the user input. Based on the user input and the input retrieval information, the question-answering big language model outputs the user output corresponding to the user input.

6. The question-answering system construction method according to claim 5, characterized in that, After constructing a question-answering engine using the aforementioned generative large language model, understanding user input based on the question-answering engine and the port and shipping database, and outputting the corresponding user output, the process further includes: Obtain negative user feedback and the corresponding interaction data; The negative user feedback, the interaction data, and the corresponding data entered into the database are manually reviewed, and the corresponding data is re-entered into the port and shipping database after the review is approved.

7. The question-answering system construction method according to claim 5, characterized in that, After constructing a question-answering engine using the aforementioned generative large language model, understanding user input based on the question-answering engine and the port and shipping database, and outputting the corresponding user output, the process further includes: Detect and identify updated data, which is data that needs to be updated into the port and shipping database; Based on the review module and the review retrieval information, the updated data is reviewed using multiple review language models. If all multiple review language models pass the review, the updated data is entered into the port and shipping database; or Based on the error identification and completion module and the completion retrieval information, the completion language model is used to identify and correct errors in the updated data and / or complete correct information in the updated data, and the corrected and / or completed updated data is entered into the port and shipping database.

8. The question-answering system construction method according to claim 7, characterized in that, After the updated data is entered into the port and shipping database, the following steps are also included: Record the time and content of the updated data being entered into the port and shipping database.

9. The question-answering system construction method according to claim 1, characterized in that, The process of acquiring port and shipping data and generating data for storage after preprocessing the port and shipping data includes: The port and shipping data is automatically collected through multiple channels, including one or more of the following: port and shipping professional literature, databases, websites, and internal company documents. The port and shipping data is deduplicated, completed, and formatted to ensure its integrity and consistency. Based on the domain and theme of the port and shipping data, the port and shipping data is classified and organized to generate the data to be stored in the database.

10. A question-answering system based on a large language model, characterized in that, include: A knowledge base construction module is used to acquire port and shipping data, port and shipping industry knowledge and industry standards, and to preprocess the port and shipping data to generate data to be stored in the database. The knowledge base review module, based on the port and shipping industry knowledge and the port and shipping industry standards, uses a generative large language model to generate review rules. The review rules work by combining retrieval enhancement generation with prompt words. The prompt words are used to guide the generation of the large language model to generate review rules, so as to help the large language model clarify the direction of the generated content based on port and shipping industry knowledge and standards, and ensure that the generated review rules can effectively cover industry standards and knowledge in the port and shipping field. Based on the review rules, the large language model is used to generate a review module. The review module includes review knowledge text and preset review prompt words. The review module can perform retrieval enhancement generation on the data in the database based at least on the port and shipping industry knowledge, the port and shipping industry standards and the review knowledge text to obtain review retrieval information corresponding to the data in the database. The audited retrieval information conforms to the port and shipping industry knowledge and the port and shipping industry standards. The preset audit prompt words are used to guide the auditing language model to audit the data in the database when the audit module is applied to the auditing language model. Based on the review module and the review retrieval information, multiple review language models are used to review the data entering the database. When all multiple review language models pass the review, the data entering the database is entered into the port and shipping database. The review module sends the data entering the database and the corresponding review retrieval information to the review language model. The review language model understands the data entering the database based on the review retrieval information and, guided by the preset review prompt words, reviews the data entering the database in combination with the review knowledge text, the review retrieval information, and its own retrieval data. The question-answering engine building module is used to build a question-answering engine using the generated large language model, and to understand user input based on the question-answering engine and the port and shipping database, and output the corresponding user output.

Citation Information

Patent Citations

  • A name list auditing system and an auditing method thereof

    CN109685463A

  • Single disease data filling and checking method, device, equipment and medium

    CN114864030A

  • Real estate data exchange system based on ownership survey data

    CN116266170A

  • Knowledge question-answering method and system based on large language model

    CN117708282A

  • Insurance claim settlement material auditing method and device based on large language model

    CN118350773A