Agricultural question and answer method, system and equipment based on multiple modes and medium

By collecting, preprocessing and vectorizing multimodal knowledge in the agricultural field, the problem that traditional question-and-answer systems cannot effectively deal with multimodal knowledge is solved, and efficient and accurate agricultural question-and-answer services and knowledge retrieval is achieved.

CN120011519APending Publication Date: 2025-05-16SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510157901.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional agricultural Q&A systems cannot effectively process and utilize multimodal knowledge, such as video, audio, pictures and text, resulting in a single question-answer content and mode.

Method used

Multimodal knowledge in the agricultural field is collected through multiple channels in the Internet, databases and literature, and knowledge crawling processing, preprocessing and vectorization processing are carried out, vector representations are generated and stored in vector databases, realizing intelligent question-and-answer and knowledge recall.

Benefits of technology

It provides integration and processing capabilities of multimodal knowledge, improves the accuracy and efficiency of question-and-answer questions and answers, optimizes user experience, and supports the popularization of agricultural knowledge and technological innovation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011519A_ABST
    Figure CN120011519A_ABST
Patent Text Reader

Abstract

The invention discloses an agricultural question-answering method, system and device based on multiple modes and a medium, belongs to the technical field of artificial intelligence, and aims to solve the technical problem that a traditional question-answering system in the agricultural field is single in question-answering content and mode. Knowledge preprocessing: video knowledge, audio knowledge, picture knowledge and text knowledge are preprocessed according to different characteristics of knowledge modalities, and the preprocessed knowledge is put in storage; according to the video knowledge, key frames are extracted through an ffmpeg video analysis technology, and key frame picture information meanings are extracted through image recognition, namely text description; audio knowledge converts audios into characters through a voice recognition technology, and the characters are cleaned and sorted into knowledge; key information is extracted from picture knowledge through an image recognition technology; performing natural language processing of word segmentation and part-of-speech tagging on the text; performing knowledge vectorization; and recalling knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal agricultural question-answering method, system, device and medium. Background Art

[0002] With the development of artificial intelligence technology, question-answering systems have been widely used in various fields. However, most existing question-answering systems can only process single-modal knowledge, such as text, and it is difficult to meet the demand for multimodal knowledge in the agricultural field. Knowledge in the agricultural field often exists in multiple forms such as video, audio, pictures and text, and traditional question-answering systems cannot effectively process and utilize this multimodal knowledge. Summary of the invention

[0003] The technical task of the present invention is to provide a multimodal agricultural question-answering method, system, equipment and medium to solve the problem of single question-answering content and mode in traditional agricultural question-answering systems.

[0004] The technical task of the present invention is achieved in the following way: a multimodal agricultural question-answering method, which is specifically as follows:

[0005] Knowledge collection: Collect multimodal knowledge in the agricultural field through the Internet, databases and documents, and clean and store the knowledge through knowledge crawlers; multimodal knowledge in the agricultural field includes agricultural knowledge, papers, books and documents;

[0006] Knowledge preprocessing: Preprocess video knowledge, audio knowledge, image knowledge and text knowledge according to the different characteristics of knowledge modalities, and store the preprocessed knowledge in the database; among them, video knowledge uses ffmpeg video analysis technology to extract key frames, and uses image recognition to extract the meaning of key frame image information, that is, text description; audio knowledge uses speech recognition technology to convert audio into text, and the text is cleaned and organized into knowledge; image knowledge uses image recognition technology to extract key information; text is processed by natural language processing such as word segmentation and part-of-speech tagging;

[0007] Knowledge vectorization: The preprocessed knowledge content is vectorized through the word2vec deep learning model to generate vector representation and store it in the vector database;

[0008] Intelligent question answering: After receiving questions from users, the system uses natural language understanding technology to analyze questions and generate query requests;

[0009] Knowledge Recall: Based on the query request, the vector database retrieves the most relevant vector representation, and uses the most relevant vector representation to recall the original knowledge form, and presents the recalled knowledge to the user in the form of text, pictures and videos.

[0010] Preferably, the knowledge processed by the knowledge crawler includes documents and structured information;

[0011] Among them, files are stored in object storage;

[0012] The structured information is stored in the structured database MYSQL.

[0013] Preferably, the structured information includes the file name, the file storage address URL, and the knowledge type and knowledge classification of the file itself.

[0014] Preferably, when knowledge is stored in the database, the source of the knowledge should be clearly identified to facilitate subsequent retrieval of multimodal resources. The format is as follows:

[0015] {“knowledge_title”:“knowledge name”, “content”:“original knowledge content”, “create_time”:“knowledge creation time”, “penetrate_data”:“other services that need to be transparently transmitted”}.

[0016] A multimodal agricultural question-answering system, the system comprising:

[0017] The knowledge collection module is used to collect multimodal knowledge in the agricultural field through the Internet, databases and documents, and clean and store the knowledge through knowledge crawlers; the multimodal knowledge in the agricultural field includes agricultural knowledge, papers, books and documents;

[0018] The knowledge preprocessing module is used to preprocess video knowledge, audio knowledge, image knowledge and text knowledge according to different characteristics of knowledge modalities, and store the preprocessed knowledge into the database;

[0019] The knowledge vectorization module is used to vectorize the preprocessed knowledge content through the word2vec deep learning model to generate vector representation and store it in the vector database; wherein the vector database is used to store the vector representation;

[0020] The intelligent question-answering module is used to obtain questions from users, analyze them through natural language understanding technology, and generate query requests;

[0021] The knowledge recall module is used to retrieve the most relevant vector representation from the vector database according to the query request, and use the most relevant vector representation to recall the original knowledge form, and present the recalled knowledge to the user in the form of text, pictures and videos.

[0022] Preferably, the knowledge processed by the knowledge acquisition module through the knowledge crawler includes files and structured information;

[0023] Among them, files are stored in object storage;

[0024] The structured information is stored in the structured database MYSQL; the structured information includes the file name, the file storage address URL, and the knowledge type and knowledge classification of the file itself.

[0025] Preferably, the knowledge preprocessing module includes:

[0026] The video analysis submodule is used to extract key frames through ffmpeg video analysis technology and extract the meaning of key frame image information, i.e. text description, through image recognition;

[0027] The audio recognition submodule is used to convert audio into text through speech recognition technology, and clean and organize the text into knowledge;

[0028] Image recognition submodule, used to extract key information through image recognition technology;

[0029] The text processing submodule is used for natural language processing such as word segmentation and part-of-speech tagging of text.

[0030] Preferably, when knowledge is stored in the database, the source of the knowledge should be clearly identified to facilitate subsequent retrieval of multimodal resources. The format is as follows:

[0031] {“knowledge_title”:“knowledge name”, “content”:“original knowledge content”, “create_time”:“knowledge creation time”, “penetrate_data”:“other services that need to be transparently transmitted”}.

[0032] An electronic device comprising: a memory and at least one processor;

[0033] Wherein, the memory stores computer-executable instructions;

[0034] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the multimodal-based agricultural question-answering method as described above.

[0035] A computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the multimodal-based agricultural question-answering method as described above is implemented.

[0036] The multimodal agricultural question-answering method, system, device and medium of the present invention have the following advantages:

[0037] (I) The present invention provides users with accurate, rich and efficient agricultural question-and-answer services, as follows:

[0038] ① Multimodal knowledge integration: It can integrate and process knowledge content in multiple modes such as video, audio, pictures and text to provide users with more comprehensive and richer information;

[0039] ② Improve the accuracy of question and answer: Through multimodal knowledge preprocessing and vectorization processing, it can more accurately understand the user's questions and provide more accurate answers;

[0040] ③ Improve knowledge retrieval efficiency: Using vector database and vector recall technology, relevant knowledge can be quickly retrieved and recalled, significantly improving the efficiency of knowledge retrieval;

[0041] ④ Optimize user experience: Intelligently present knowledge and provide personalized services based on users’ questioning habits and preferences, thereby optimizing user experience;

[0042] ⑤ Support the popularization of agricultural knowledge: It can be widely used in agricultural education and training, helping farmers and agricultural workers to quickly acquire agricultural knowledge and promote the popularization and dissemination of agricultural knowledge;

[0043] ⑥ Promote agricultural technology innovation: By providing rich agricultural knowledge resources and efficient question-and-answer services, it can stimulate technological innovation and research in the agricultural field and promote the development of agricultural science and technology;

[0044] ⑦ Enhanced agricultural decision support: It can provide agricultural managers and decision makers with decision support based on big data, helping them make more scientific and reasonable decisions;

[0045] ⑧ Lower the threshold for knowledge acquisition: Through intelligent question-and-answer services, the threshold for users to acquire agricultural knowledge can be lowered, allowing non-professionals to easily obtain the required information;

[0046] ⑨Promote agricultural informatization: The application helps to promote the process of agricultural informatization and improve the level of intelligent agricultural production and management;

[0047] (ii) The present invention combines multimodal knowledge recall technology to provide more comprehensive, accurate and interactive agricultural information consulting services, which not only supports traditional text-based knowledge, but also innovatively incorporates multimodal knowledge types such as video, audio, and pictures. Through advanced preprocessing technology, these diversified knowledge contents are converted into digital information that is easy to store, retrieve and apply;

[0048] (III) In the knowledge preprocessing stage, the present invention fully considers the uniqueness and complexity of different modal knowledge; for video, video analysis technology is used to extract key frames, and the key frame information is subjected to feature extraction and converted into textual knowledge description; for audio, speech recognition and audio analysis technology are used to convert it into text or labels; for pictures, key features and description information are extracted through image recognition technology; and textual knowledge is subjected to semantic analysis and knowledge extraction through natural language processing technology, so as to retain the core value and information integrity of the original knowledge to the greatest extent;

[0049] (iv) After preprocessing, the multimodal knowledge is converted into a unified digital expression form through vectorization technology and stored in an efficient vector database, which not only improves the storage efficiency of knowledge, but also provides a solid foundation for subsequent intelligent question answering and knowledge association;

[0050] (V) In the intelligent question-answering phase, the present invention can recall multimodal knowledge related to the question from the vector database according to the user's question; at the same time, the original knowledge form (such as video, audio, picture, etc.) is accurately recalled through the association relationship established when entering the database, presenting rich and diverse answers and explanations to the user. This cross-modal knowledge fusion and association capability has significant advantages and potential in the intelligent question-answering application in the agricultural field;

[0051] (VI) The present invention can make full use of multimodal knowledge to improve the accuracy and richness of questions and answers. At the same time, through vectorization processing and the use of vector databases, the efficiency of knowledge retrieval and recall is improved; at the same time, it can process knowledge content in multiple modes including video, audio, pictures and text, and can support accurate recall when answering questions, realizing the rich multimodal display of answers, solving the problem of the single answer content mode of traditional question-and-answer systems in the agricultural field, and can efficiently process and answer agricultural-related questions. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The present invention is further described below in conjunction with the accompanying drawings.

[0053] Attached Figure 1 It is a flowchart of the agricultural question answering method based on multimodality;

[0054] Attached Figure 2 A flowchart for knowledge collection;

[0055] Attached Figure 3 A flowchart for audio knowledge preprocessing;

[0056] Attached Figure 4 A flowchart for video knowledge preprocessing;

[0057] Attached Figure 5 A flowchart for image knowledge preprocessing;

[0058] Attached Figure 6 A flowchart for knowledge vectorization;

[0059] Attached Figure 7 This is a flowchart of the intelligent question answering process;

[0060] Attached Figure 8 A flowchart of the knowledge recall process. DETAILED DESCRIPTION

[0061] The multimodal agricultural question-and-answer method, system, device and medium of the present invention are described in detail below with reference to the drawings and specific embodiments of the specification.

[0062] Embodiment 1:

[0063] As attached Figure 1 As shown, this embodiment provides an agricultural question-answering method based on multimodality, and the method is specifically as follows:

[0064] S1. Knowledge collection: as attached Figure 2 As shown, multimodal knowledge in the agricultural field is collected through multiple channels such as the Internet, databases and documents, and the knowledge is cleaned and stored through knowledge crawler processing; among them, multimodal knowledge in the agricultural field includes agricultural knowledge, papers, books and documents;

[0065] S2. Knowledge preprocessing: Preprocess video knowledge, audio knowledge, image knowledge and text knowledge according to the different characteristics of knowledge modalities, and store the preprocessed knowledge in the database; Figure 4 As shown in the figure, the video knowledge extracts key frames through ffmpeg video analysis technology, and extracts the meaning of key frame picture information through image recognition, that is, text description; as shown in the attached Figure 3 As shown in the figure, audio knowledge is converted into text through speech recognition technology, and the text is cleaned and organized into knowledge; as shown in the attached Figure 5 As shown in the figure, image knowledge extracts key information through image recognition technology; text is processed by natural language segmentation and part-of-speech tagging;

[0066] S3. Knowledge vectorization: as shown in the attached Figure 6 As shown, the preprocessed knowledge content is vectorized through the word2vec deep learning model to generate a vector representation and store it in a vector database;

[0067] S4. Intelligent Question and Answer: As shown in the attached Figure 7 As shown, after obtaining the question raised by the user, the question is parsed through natural language understanding technology and a query request is generated;

[0068] S5. Knowledge recall: as attached Figure 8As shown, according to the query request, the vector database retrieves the most relevant vector representation, and uses the most relevant vector representation to recall the original knowledge form, and presents the recalled knowledge to the user in the form of text, pictures and videos.

[0069] The knowledge processed by the knowledge crawler in step S1 of this embodiment includes files and structured information;

[0070] Among them, files are stored in object storage;

[0071] The structured information is stored in the structured database MYSQL.

[0072] The structured information in this embodiment includes the file name, the file storage address URL, and the knowledge type and knowledge classification of the file itself.

[0073] When the knowledge in step S2 of this embodiment is stored in the database, the source of the knowledge is clarified to facilitate the subsequent recall of multimodal resources. The format is as follows:

[0074] {“knowledge_title”:“knowledge name”, “content”:“original knowledge content”, “create_time”:“knowledge creation time”, “penetrate_data”:“other services that need to be transparently transmitted”}.

[0075] Embodiment 2:

[0076] This embodiment provides a multimodal agricultural question-answering system, which includes:

[0077] The knowledge collection module is used to collect multimodal knowledge in the agricultural field through the Internet, databases and documents, and clean and store the knowledge through knowledge crawlers; the multimodal knowledge in the agricultural field includes agricultural knowledge, papers, books and documents;

[0078] The knowledge preprocessing module is used to preprocess video knowledge, audio knowledge, image knowledge and text knowledge according to different characteristics of knowledge modalities, and store the preprocessed knowledge into the database;

[0079] The knowledge vectorization module is used to vectorize the preprocessed knowledge content through the word2vec deep learning model to generate vector representation and store it in the vector database; wherein the vector database is used to store the vector representation;

[0080] The intelligent question-answering module is used to obtain questions from users, analyze them through natural language understanding technology, and generate query requests;

[0081] The knowledge recall module is used to retrieve the most relevant vector representation from the vector database according to the query request, and use the most relevant vector representation to recall the original knowledge form, and present the recalled knowledge to the user in the form of text, pictures and videos.

[0082] The knowledge processed by the knowledge acquisition module in this embodiment through the knowledge crawler includes files and structured information;

[0083] Among them, files are stored in object storage;

[0084] The structured information is stored in the structured database MYSQL; the structured information includes the file name, the file storage address URL, and the knowledge type and knowledge classification of the file itself.

[0085] The knowledge preprocessing module in this embodiment includes:

[0086] The video analysis submodule is used to extract key frames through ffmpeg video analysis technology and extract the meaning of key frame image information, i.e. text description, through image recognition;

[0087] The audio recognition submodule is used to convert audio into text through speech recognition technology, and clean and organize the text into knowledge;

[0088] Image recognition submodule, used to extract key information through image recognition technology;

[0089] The text processing submodule is used for natural language processing such as word segmentation and part-of-speech tagging of text.

[0090] In this embodiment, when knowledge is stored, the source of the knowledge is clarified to facilitate the subsequent recall of multimodal resources. The format is as follows:

[0091] {“knowledge_title”:“knowledge name”, “content”:“original knowledge content”, “create_time”:“knowledge creation time”, “penetrate_data”:“other services that need to be transparently transmitted”}.

[0092] Embodiment 3:

[0093] This embodiment also provides an electronic device, including: a memory and at least one processor;

[0094] Wherein, the memory stores computer-executable instructions;

[0095] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the multimodal-based agricultural question-answering method described in any one of the present inventions.

[0096] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.

[0097] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state storage devices.

[0098] Embodiment 4:

[0099] This embodiment also provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions are loaded by a processor, so that the processor executes the agricultural question-answering method based on multimodality in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code that implements the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0100] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.

[0101] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.

[0102] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.

[0103] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multimodal agricultural question-answering method, characterized in that: The method is as follows: Knowledge collection: Collect multimodal knowledge in the agricultural field through the Internet, databases and literature, and clean and store the knowledge through knowledge crawlers; Knowledge preprocessing: Preprocess video knowledge, audio knowledge, image knowledge and text knowledge according to the different characteristics of knowledge modalities, and store the preprocessed knowledge in the database; among them, video knowledge uses ffmpeg video analysis technology to extract key frames, and uses image recognition to extract the meaning of key frame image information, that is, text description; audio knowledge uses speech recognition technology to convert audio into text, and the text is cleaned and organized into knowledge; image knowledge uses image recognition technology to extract key information; text is processed by natural language processing such as word segmentation and part-of-speech tagging; Knowledge vectorization: The preprocessed knowledge content is vectorized through the word2vec deep learning model to generate vector representation and store it in the vector database; Intelligent question answering: After receiving questions from users, the system uses natural language understanding technology to analyze questions and generate query requests; Knowledge Recall: Based on the query request, the vector database retrieves the most relevant vector representation, and uses the most relevant vector representation to recall the original knowledge form, and presents the recalled knowledge to the user in the form of text, pictures and videos.

2. The multimodal agricultural question-answering method according to claim 1, characterized in that: The knowledge processed by knowledge crawlers includes documents and structured information; Among them, files are stored in object storage; The structured information is stored in the structured database MYSQL.

3. The multimodal agricultural question-answering method according to claim 2, characterized in that: The structured information includes the file name, the file storage address URL, and the knowledge type and knowledge classification of the file itself.

4. The multimodal agricultural question-answering method according to claim 3, characterized in that: When knowledge is stored in the database, the format is as follows: {"knowledge_title":"Knowledge name","content":"Original knowledge content to be transmitted","create_time":"Knowledge creation time","penetrate_data":"Other services that require transparent transmission"}.

5. A multimodal agricultural question-answering system, characterized in that: The system includes: The knowledge acquisition module is used to collect multimodal knowledge in the agricultural field through the Internet, databases and literature, and clean and store the knowledge through knowledge crawler processing; The knowledge preprocessing module is used to preprocess video knowledge, audio knowledge, image knowledge and text knowledge according to different characteristics of knowledge modalities, and store the preprocessed knowledge into the database; The knowledge vectorization module is used to vectorize the preprocessed knowledge content through the word2vec deep learning model to generate vector representation and store it in the vector database; wherein the vector database is used to store the vector representation; The intelligent question-answering module is used to obtain questions from users, analyze them through natural language understanding technology, and generate query requests; The knowledge recall module is used to retrieve the most relevant vector representation from the vector database according to the query request, and use the most relevant vector representation to recall the original knowledge form, and present the recalled knowledge to the user in the form of text, pictures and videos.

6. The multimodal agricultural question-answering system according to claim 5, characterized in that: The knowledge collected by the knowledge acquisition module through the knowledge crawler includes documents and structured information; Among them, files are stored in object storage; The structured information is stored in the structured database MYSQL; the structured information includes the file name, the file storage address URL, and the knowledge type and knowledge classification of the file itself.

7. The multimodal agricultural question-answering system according to claim 5, characterized in that: The knowledge preprocessing module includes: The video analysis submodule is used to extract key frames through ffmpeg video analysis technology and extract the meaning of key frame image information, i.e. text description, through image recognition; The audio recognition submodule is used to convert audio into text through speech recognition technology, and clean and organize the text into knowledge; Image recognition submodule, used to extract key information through image recognition technology; The text processing submodule is used for natural language processing such as word segmentation and part-of-speech tagging of text.

8. The multimodal agricultural question-answering system according to any one of claims 5 to 7, characterized in that: When knowledge is stored in the database, the format is as follows: {"knowledge_title":"Knowledge name","content":"Original knowledge content to be transmitted","create_time":"Knowledge creation time","penetrate_data":"Other services that require transparent transmission"}.

9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the multimodal-based agricultural question-answering method as described in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the multimodal agricultural question-answering method as described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Agricultural intelligent question and answer method and system supporting multi-mode and multi-round dialogues

    CN117370506A

  • Agricultural multi-mode intelligent retrieval technology and system based on multi-source heterogeneous data

    CN117573882A

  • Large language model knowledge question-answering method and system fused with multi-modal knowledge graph

    CN118627628A