A Construction Method of a Language Model for the Forestry Vertical Domain

By improving the ViLBERT model and building the RAG knowledge base, the problem of insufficient image and text matching accuracy in forestry management is solved, and the high quality and flexibility of forestry intelligent Q&A and document generation is achieved.

CN120068947BActive Publication Date: 2025-07-22JIANGXI WOODPECKER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510546350.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-22
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In the prior art, the matching accuracy between images and text in forestry management is insufficient, resulting in inaccurate questions and answers generated in intelligent questions and answers.

Method used

Construct a language model for the vertical field of forestry. By improving the ViLBERT model, we introduce learnable weights in the Co-TRM layer, combine multimodal coding, weighted similarity calculation and frozen text encoder training LLM, and build a RAG knowledge base for intelligent Q&A.

Benefits of technology

It significantly improves the matching accuracy of images and text and the accuracy of information understanding, ensures the high quality and document generation of forestry intelligent question-and-answer system, and provides flexible retrieval and accurate text generation capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068947B_ABST
    Figure CN120068947B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a large language model for the forestry vertical field, belonging to the technical field of computer models, including constructing a first dataset for forestry management business Q&A; constructing a multimodal encoding model; freezing the text stream of the multimodal encoding model, and forming a large language model by combining the multimodal encoding model and the LLM decoder and training it to generate target text related to the user input when generating intelligent Q&A; constructing a RAG knowledge base; and performing intelligent Q&A based on RAG retrieval. In view of the highly matching characteristics of forestry management business graphics and texts, the present invention introduces innovative technologies such as multimodal joint encoding, weighted similarity calculation, and freezing the text encoder to train the LLM, greatly improving the quality and application value of forestry intelligent Q&A and document generation. It also solves the problem of insufficient matching accuracy between existing images and texts, and can achieve more flexible and accurate retrieval and intelligent Q&A in combination with RAG technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer models, and particularly to a method for constructing a large language model for the forestry vertical domain. Background Art

[0002] In forestry management, especially in the business process of "forest right acceptance, review, and certification", the accurate acquisition of data and information is crucial. With the rapid development of image and text fusion technology, large language models based on multimodal learning have shown great potential in the fields of forestry intelligent question answering, business process handling, and document generation. As a multimodal deep learning method, ViLBERT not only significantly improves the accuracy of information understanding but also ensures the matching accuracy of images and texts by introducing joint encoding of images and texts, thus promoting the development of forestry intelligent systems.

[0003] In forestry business processing and intelligent question answering, ensuring the efficient fusion of graphic and text information is crucial for performing accurate question answering and document generation. Related application requirements often require the close combination of images and texts to meet the needs of business process and knowledge question answering. However, the limitations of existing technologies result in insufficient matching accuracy between images and texts, leading to inaccurate question answering in intelligent question answering. In forestry business management, image data (such as satellite images of land or forest areas) and text data (such as forest right documents) are closely related. How to use this feature to achieve high-precision matching of images and texts and apply it to intelligent question answering is a problem that needs to be solved.

[0004] ViLBERT (Vision and Language BERT) is a multi-modal pre-trained model designed to process both image and text data simultaneously. It consists of two parallel BERT (Bidirectional Encoder Representation from Transformer) streams, an image stream and a text stream, for processing image data and text data respectively. Each stream is composed of multiple TRM (Transformer) and Co-TRM layers (co-attentional TRM layers). The biggest improvement of Co-TRM lies in swapping the K matrix and V matrix of the two streams, thus achieving information fusion between the two different streams. Therefore, the number of Co-TRM layers in the image stream and the text stream should be the same, and they perform information fusion in a one-to-one correspondence. For example, if there are N Co-TRM layers in each stream, the nth Co-TRM layer in the image stream should correspond one-to-one with the nth Co-TRM layer in the text stream. In the corresponding two Co-TRM layers, the Q matrix of the Co-TRM layer in the image stream comes from image information, while the K matrix and V matrix come from text information. The Q matrix of the Co-TRM layer in the text stream comes from text information, and the K matrix and V matrix come from image information.

[0005] ViLBERT uses two pre-training tasks:

[0006] Task 1: Masked multi-modal modelling task, which includes MLM and MRM; MLM (Masked Language Model), the task is to predict some randomly masked words in the input sentence. MRM (Masked Region Model), randomly masks some regions in the image, and uses the remaining part of the image to predict the content of the masked region.

[0007] Task 2: Multi-modal alignment prediction task, the model presents image-text pairs and predicts whether the image and text are aligned.

[0008] A large language model (LLM) refers to a deep learning model trained using a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text. Its principle is as follows: First, the input layer receives the text data input by the user and converts it into a vector representation form that the model can process, that is, performs operations such as word embedding to map words or sentences into a low-dimensional vector space. Then, the encoder based on the Transformer architecture extracts semantic information and context information in the text. The decoder, according to the features extracted by the encoder and the previous generated text information, gradually generates the next word or character as output. By predicting the next possible word or token, a complete text content is continuously constructed to achieve tasks such as language generation and question answering. Summary of the Invention

[0009] The object of the present invention is to provide a construction method of a language large model for the forestry vertical field, which solves the problem of insufficient accuracy of existing image and text matching in the fields of forestry management and resource approval, resulting in inaccurate questions and answers generated in intelligent question answering.

[0010] To achieve the above object, the technical solution adopted by the present invention is as follows: A construction method of a language large model for the forestry vertical field includes the following steps;

[0011] S1, construct the first dataset D1 for forestry management business Q&A;

[0012] Determine multiple types of forestry management businesses. For each type of business, obtain multiple texts and the corresponding images for each text, and form text-image pairs from the corresponding texts and images. All text-image pairs form the first dataset D1;

[0013] S2, construct a multi-modal encoding model, including S21~S22;

[0014] S21, construct an improved ViLBERT model;

[0015] Obtain a ViLBERT model, whose image stream and text stream each include N Co-TRM layers. For the nth Co-TRM layer in the image stream and the nth Co-TRM layer in the text stream , construct the first improved Q matrix according to the following formula Replace the Q matrix Q1 in, the second improved Q matrix Replace the Q matrix Q2 in to obtain an improved ViLBERT model;

[0016] ,

[0017] ,

[0018] In the formula, , are respectively the learnable weights of Q1 and Q2 in , are respectively the learnable weights of Q2 and Q1 in, 1 ≤ n ≤ N;

[0019] S22. Use D1 to train and improve the ViLBERT model based on the MLM task, MRM task, and multi-modal alignment task until convergence to obtain a multi-modal encoding model. The multi-modal model is used to input image-text pairs, and the image and text in the image-text pair are respectively output as image encoding and text encoding through the image stream and text stream;

[0020] S3. Train a large language model;

[0021] Freeze the text stream of the multi-modal encoding model, and form a large language model with the multi-modal encoding model and the LLM decoder, and train the large language model to generate a text vector v according to the user input q , and send it to the LLM decoder to generate the target text related to the user input when generating business intelligent answers;

[0022] If the user input is text, the corresponding text encoding obtained by the multi-modal model is used as v q , if the user input is an image, the multi-modal model performs a multi-modal alignment task and outputs the text encoding of the text aligned with the image as v q ;

[0023] S4. Construct a RAG knowledge base D2. The k-th data in D2 is , where T k , I k are respectively the text and image in the k-th image-text pair in D1, , are respectively the text encoding and image encoding obtained by T k , I k through the multi-modal encoding model;

[0024] S5. Perform intelligent answers based on RAG retrieval;

[0025] S51. Input the user's query data into the multi-modal encoding model to obtain the corresponding encoding. The query data includes a query image and / or query text, and the corresponding encoding is the query image encoding and / or query text encoding ;

[0026] S52. Calculate the score and weight of each data in D2. Among them, the k-th data dk Score and weight are obtained according to the following formula;

[0027] ,

[0028] ,

[0029] where γ is the score weight, is the calculation of the modulus length, is the exp function;

[0030] S53. For the preset threshold, the data with scores greater than the threshold are screened out and placed in the candidate set B, and the data in the candidate set are weighted according to the following formula to obtain a weighted result T RAG ;

[0031] ,

[0032] S54. Generate a prompt vector based on the query data, and based on T RAG and the prompt vector, the target text is generated by the LLM decoder.

[0033] Preferably: In S1, the service is forest right acceptance, review and certification, the text is the text data corresponding to various services, and the image is the remote sensing image and satellite image corresponding to the text.

[0034] Preferably: In S21, constructing an improved ViLBERT model is specifically;

[0035] Obtain a ViLBERT model including an image stream and a text stream;

[0036] The image stream includes an image encoding layer, N Co-TRM layers, and a first TRM layer arranged in sequence. The nth Co-TRM layer in the image stream is ;

[0037] The text stream includes a text encoding layer, a second TRM layer, N Co-TRM layers, and a third TRM layer arranged in sequence. The nth Co-TRM layer in the text stream is ;

[0038] is a Transformer layer, including a Q matrix Q1, a K matrix, a V matrix, a multi-head attention layer, a multi-layer perceptron, and a batch normalization layer. The Q1 is from the image encoding layer, and the K matrix and the V matrix are from the second TRM layer, and are obtained according to the formula to obtain , and replace Q1 and send it into the multi-head attention layer;

[0039] It is a Transformer layer, including a Q matrix Q2, a K matrix, and a V matrix. The Q matrix is from the second TRM layer, and the K matrix and the V matrix are from the image encoding layer. According to the formula to obtain , which is sent to the multi-head attention layer instead of Q2.

[0040] Preferably, when training the improved ViLBERT model in S22, the total loss L total is used to adjust the model parameters, where;

[0041] ,

[0042] ,

[0043] ,

[0044] ,

[0045] ,

[0046] ,

[0047] L total In L MLM , , L matching are the loss functions of the MLM task, the MRM task, and the multi-modal alignment task respectively. α1, α2, and α3 are the weights of L MLM , , and L matching respectively;

[0048] In L MLM , M T is the set of positions of the masked words in the text, t i is the true label of the i-th masked word, T i is the input sequence obtained by masking the text, and P(t i │T i ) is the probability that the i-th masked word is predicted as t i ;

[0049] In , are the input image and the output image of the improved ViLBERT model when performing the MRM task respectively. The input image is an image of a randomly occluded partial area in a text-image pair, and the output image is an image predicted for the occluded area based on the occluded area and the text in the text-image pair. is the L2 norm;

[0050] Lmatching Among them, is the sigmoid function, and v T , v I are respectively the text encoding and image encoding obtained by an improved ViLBERT model for a text-image pair in D1;

[0051] is the similarity function for calculating similarity, w(v T , v I ) is the dynamic weighting factor, is the exp function, and β is the adjustment factor.

[0052] Preferably: When training the large language model in S3, text data is generated by the autoregressive method, and the loss function is L generation ;

[0053] ,

[0054] In the formula, where: t z is the z-th target word in the target text during the generation process, Z is the total number of target words in the target text, is the probability of generating t q based on the generated text and v z , and L generation trains the large language model by maximizing the conditional probability of the target text.

[0055] Preferably: S54 is specifically:

[0056] If the query data is a query text, the prompt vector is , if the query data is a query image, then the prompt vector is , take T RAG as the K matrix and V matrix in the LLM decoder, and the prompt vector as the Q matrix in the LLM decoder, and the target text is generated by the LLM decoder;

[0057] If the query data is a query image and a query text, first set the prompt vector to , and generate a text Text1 with T RAG ; Concatenate Text1 with the query text, and then obtain a text encoding through the text stream of the multimodal model, set the prompt vector to , and generate the target text with T RAG .

[0058] Compared with the prior art, the advantages of the present invention are:

[0059] (1) According to the characteristic that in forestry management, image data (such as satellite images of land or forest areas) is closely related to text data (such as certain types of documents), multi-modal joint encoding is introduced to improve the ViLBERT model. Four learnable weights are introduced in the Co-TRM layer, which can not only significantly improve the accuracy of information understanding, but also ensure the matching accuracy between images and texts, providing stronger data support for the forestry intelligent question-answering system, and thus promoting the development of the forestry intelligent system.

[0060] (2) When training the improved ViLBERT model, a dynamic weighting factor w(v matching ,v T ) is introduced into the loss function L I of its multi-modal alignment task. When the similarity is high, w(v T ,v I ) will approach 1, indicating a high matching degree between the text and the image, and the model should quickly update the parameters. When the similarity is low, w(v T ,v I ) will approach 0, indicating a low matching degree between the text and the image, and the model should slow down the update speed. In this way, the text-image pairs with higher similarity are given higher weights, thus accelerating the training.

[0061] (3) By freezing the text stream part of the multi-modal encoding model and training a large language model based on the output of the text stream for text generation, the consistency and business relevance of text generation are ensured. Especially in the generation of forestry documents, this method provides an innovative technical path for high-quality text output that meets actual needs.

[0062] (4) In the RAG retrieval of the present invention, multiple results are weighted and fused, and then a prompt vector is generated according to the query data, and target text generation is performed together with the weighted results of the RAG retrieval. Since a prompt vector can be obtained regardless of whether the input query data is a query text, a query image, or a combination of both, the flexibility of the retrieval is improved, and while ensuring flexibility, the accuracy of the target text is also taken into account.

[0063] In summary, the present invention has greatly improved the quality and application value of forestry intelligent question-answering and document generation through innovative technologies such as introducing multi-modal joint encoding, weighted similarity calculation, and freezing the text encoder to train the LLM, providing an innovative solution for forestry-related fields. By solving the problem of insufficient matching accuracy between existing images and texts, this method provides a new technical path for forestry business processes, intelligent question-answering, and document generation, and has broad application prospects in various businesses in the forestry management field. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is the structural diagram of the improved ViLBERT model of the present invention;

[0065] Figure 2 To improve the interaction schematic diagram in the ViLBERT model and . Specific implementation manner

[0066] The present invention will be further described below in conjunction with embodiments and drawings.

[0067] Embodiment 1: Refer to Figure 1 and Figure 2 , a method for constructing a language large model for the forestry vertical field, including the following steps;

[0068] S1, construct the first data set D1 for forestry management business Q&A;

[0069] Determine multiple types of forestry management operations. For each type of operation, obtain multiple texts and the images corresponding to each text, and form text-image pairs with the corresponding texts and images. All text-image pairs constitute the first data set D1;

[0070] S2, construct a multi-modal encoding model, including S21~S22;

[0071] S21, construct an improved ViLBERT model;

[0072] Obtain a ViLBERT model, whose image stream and text stream each include N Co-TRM layers. For the nth Co-TRM layer in the image stream and the nth Co-TRM layer in the text stream , construct the first improved Q matrix according to the following formula Replace the Q matrix Q1 in, the second improved Q matrix Replace the Q matrix Q2 in to obtain an improved ViLBERT model;

[0073] ,

[0074] ,

[0075] In the formula, , are respectively the learnable weights of Q1 and Q2 in, , are respectively the learnable weights of Q2 and Q1 in, 1≤n≤N;

[0076] S22. Use D1 to train the improved ViLBERT model based on the MLM task, MRM task, and multi-modal alignment task until convergence to obtain a multi-modal encoding model. The multi-modal model is used to input image-text pairs. In the image-text pair, the image and text are respectively output as image encoding and text encoding through the image stream and text stream.

[0077] S3. Train a large language model.

[0078] Freeze the text stream of the multi-modal encoding model. Combine the multi-modal encoding model and the LLM decoder to form a large language model, and train the large language model to generate a text vector v according to the user input q , and send it to the LLM decoder to generate the target text related to the user input when generating business intelligent answers.

[0079] If the user input is text, use the multi-modal model to obtain the corresponding text encoding as v q , if the user input is an image, then the multi-modal model performs the multi-modal alignment task and outputs the text encoding of the text aligned with the image as v q ;

[0080] S4. Construct a RAG knowledge base D2. The k-th data in D2 is , where T k , I k are respectively the text and image in the k-th image-text pair in D1, , are respectively the text encoding and image encoding obtained by T k , I k through the multi-modal encoding model;

[0081] S5. Perform intelligent answering based on RAG retrieval;

[0082] S51. Input the user's query data into the multi-modal encoding model to obtain the corresponding encoding. The query data includes a query image and / or query text, and the corresponding encoding is the query image encoding and / or query text encoding ;

[0083] S52. Calculate the score and weight of each data in D2. The score k and weight of the k-th data d are obtained according to the following formula;

[0084] ,

[0085] ,

[0086] In the formula, γ is the score weight, is the calculation modulus length, is the exp function;

[0087] S53, a preset threshold, filters out the data with scores greater than the threshold and places them in the candidate set B, and the data in the candidate set is weighted according to the following formula to obtain a weighted result T RAG ;

[0088] ,

[0089] S54, generates a prompt vector based on the query data, and based on T RAG and the prompt vector, the target text is generated by the LLM decoder.

[0090] In S1 of this embodiment, the service is forest right acceptance, review, and certification, the text is the text data corresponding to various services, and the image is the remote sensing image and satellite image corresponding to the text.

[0091] In S21, constructing an improved ViLBERT model is specifically as follows;

[0092] Obtain a ViLBERT model including an image stream and a text stream;

[0093] The image stream includes an image encoding layer, N Co-TRM layers, and a first TRM layer arranged in sequence. The nth Co-TRM layer in the image stream is ;

[0094] The text stream includes a text encoding layer, a second TRM layer, N Co-TRM layers, and a third TRM layer arranged in sequence. The nth Co-TRM layer in the text stream is ;

[0095] is a Transformer layer, including a Q matrix Q1, a K matrix, a V matrix, a multi-head attention layer, a multi-layer perceptron, and a batch normalization layer. The Q1 is from the image encoding layer, the K matrix and the V matrix are from the second TRM layer, and according to the formula to obtain , and substitute it for Q1 and send it into the multi-head attention layer;

[0096] is a Transformer layer, including a Q matrix Q2, a K matrix, a V matrix. The Q matrix is from the second TRM layer, the K matrix and the V matrix are from the image encoding layer, and according to the formula to obtain , and substitute it for Q2 and send it into the multi-head attention layer. In Figure 1 after substituting for Q1, the Co-TRM layer in the original image stream is called the first improved Co-TRM layer, and After replacing Q2, the Co-TRM layer in the original text stream is called the second improved Co-TRM layer.

[0097] When training the improved ViLBERT model in S22, minimize the total loss L total Adjust the model parameters, where;

[0098] ,

[0099] ,

[0100] ,

[0101] ,

[0102] ,

[0103] ,

[0104] L total In, L MLM , , L matching are the loss functions for the MLM task, MRM task, and multi-modal alignment task respectively, and α1, α2, α3 are the weights of L MLM , , L matching respectively;

[0105] L MLM In, M T is the set of positions of the masked words in the text, t i is the true label of the i-th masked word, T i is the input sequence obtained by masking the text, P(t i │T i ) is the probability that the i-th masked word is predicted as t i ;

[0106] In, , are the input image and output image of the improved ViLBERT model when performing the MRM task respectively. The input image is the image of the randomly occluded part of the text-image pair, and the output image is the image predicted for the occluded area according to the occluded area and the text in the text-image pair. is the L2 norm;

[0107] L matching In, is the sigmoid function, v T , v IThey are the text encoding and image encoding obtained by improving the ViLBERT model for a text-image pair in D1 respectively;

[0108] The similarity function for calculating similarity is similarity, and w(v T , v I ) is the dynamic weighting factor, where is the exp function and β is the adjustment factor.

[0109] When training the large language model in S3, autoregressive method is used to generate text data, and the loss function is L generation ;

[0110] ,

[0111] In the formula, where: t z is the z-th target word in the target text during the generation process, Z is the total number of target words in the target text, is the probability of generating t q calculated based on the generated text and v z , and L generation trains the large language model by maximizing the conditional probability of the target text.

[0112] Specifically, S54 is as follows: If the query data is a query text, the prompt vector is , if the query data is a query image, then the prompt vector is , take T RAG as the K matrix and V matrix in the LLM decoder, and the prompt vector as the Q matrix in the LLM decoder to generate the target text by the LLM decoder; if the query data is a query image and a query text, first set the prompt vector to , generate a text Text1 with T RAG ; concatenate Text1 with the query text, and then obtain a text encoding through the text stream of the multimodal model, set the prompt vector to , and generate the target text with T RAG .

[0113] Regarding the training of the multimodal encoding model, it is through the following three tasks:

[0114] MLM task: In the intelligent question-answering system related to forestry, it is necessary to convert the user's query into relevant business content and consider the professionalism of forestry knowledge during this process. During the training process, through the masking mechanism, real text and image missing scenarios can be simulated, thereby improving the robustness of the model. By masking the key information in the business-related text, the model can learn to correctly fill in the missing words given the context.

[0115] MRM Task: The masked loss of the image part is very important for training the model to understand image details. For example, in forest resource management, some areas of satellite images may be missing or occluded, and the model needs to be able to infer the content of the missing parts through known areas.

[0116] Multimodal Alignment Task: In the "intelligent question answering" session, images are matched with texts to ensure the semantic consistency of the input images and corresponding texts. Through dynamic weighting, image-text pairs with higher similarity are given higher weights, thus accelerating the training.

[0117] Regarding the improved large language model: The large language model is an encoder-decoder architecture. Before training the large language model, this invention first uses a multimodal encoding model as a new encoder to replace the original encoder of the large language model, and forms an improved large language model by combining the multimodal encoding model and the LLM decoder, and then trains it according to the existing technical methods. Freezing the text stream part of the multimodal encoding model means that the parameters of the text stream will not be updated during the subsequent training process. When the multimodal encoding model inputs an image or text, it will generate a text vector according to step S3 , and the LLM decoder will generate a target text related to the query data according to .

[0118] Regarding RAG Retrieval: In various operations related to intelligent question answering in forestry management, it is necessary to efficiently retrieve a large amount of image-text data to quickly provide accurate query results for users. This invention weights and fuses multiple data obtained by RAG retrieval to generate a unified semantic representation, which is used as the input of the text generation model together with the prompt vector for the generation of the target text, and has the characteristics of flexible modal input, unified and accurate retrieval expression. Based on the RAG retrieval mechanism, this invention can perform intelligent screening according to the similarity between texts and images, thus providing accurate responses for the "intelligent question answering" system.

[0119] Regarding the LLM Decoder and Its Application: In large language models, especially in models based on the Transformer architecture, the Q matrix, K matrix, and V matrix are the core components in the self-attention mechanism. When training the LLM decoder in S3, the Q matrix, K matrix, and V matrix are directly obtained from the text vector, and the target text is finally generated. When applying the LLM decoder subsequently, the Q matrix, K matrix, and V matrix are generated according to step S5 by combining the prompt vector and T RAG to obtain the target text.

[0120] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A construction method of a language large model for the forestry vertical field, characterized in that: Including the following steps; S1. Construct the first dataset D1 of forestry management business Q&A; Determine multiple types of forestry management business. For each type of business, obtain multiple texts and the corresponding images for each text, form text-image pairs with the corresponding texts and images, and all text-image pairs form the first dataset D1; S2. Construct a multi-modal encoding model, including S21~S22; S21. Construct an improved ViLBERT model; Obtain a ViLBERT model, where both the image stream and the text stream include N Co-TRM layers, and for the nth Co-TRM layer in the image stream and the nth Co-TRM layer in the text stream , construct a first improved Q matrix according to the following formula Replace the Q matrix Q1 in and the second improved Q matrix replace the Q matrix Q2 in to obtain an improved ViLBERT model; , , In the formula, and are respectively the learnable weights of Q1 and Q2 in and are respectively the learnable weights of Q2 and Q1 in, 1 ≤ n ≤ N; S22. Use D1 to train the improved ViLBERT model based on the MLM task, MRM task, and multi-modal alignment task until convergence to obtain a multi-modal encoding model. The multi-modal encoding model is used to input text-image pairs, and the image and text in the text-image pair are respectively output as image encoding and text encoding through the image stream and text stream; S3. Train a large language model; Freeze the text stream of the multimodal encoding model, combine the multimodal encoding model and the LLM decoder to form a large language model, and train the large language model to generate a text vector v according to the user input. q Send it to the LLM decoder to generate the target text related to the user input when generating business intelligent answers. If the user input is text, the corresponding text encoding is obtained by the multimodal encoding model as v q , if the user input is an image, the multimodal encoding model performs a multimodal alignment task and outputs the text encoding of the text aligned with the image as v q ; S4, construct the RAG knowledge base D2, and the k-th data in D2 is , where T k and I k are respectively the text and image in the k-th text-image pair in D1, , are respectively the text encoding and image encoding obtained by the multi-modal encoding model for T k and I k ; S5. Conduct intelligent Q&A based on RAG retrieval; S51. Input the user's query data into the multi-modal encoding model to obtain the corresponding encoding. The query data includes a query image and / or a query text, and the corresponding encoding is the query image encoding and / or the query text encoding ; S52, calculate the score and weight of each piece of data in D2, where the score k of the k-th piece of data d and the weight are obtained according to the following formula; , , where γ is the scoring weight, is for calculating the modulus length, is the exp function; S53. A preset threshold value is used to screen out the data with scores greater than the threshold value and place them in the candidate set B. The data in the candidate set is weighted according to the following formula to obtain a weighted result T RAG ; , S54, Generate a prompt vector based on the query data, and generate the target text by the LLM decoder based on T RAG and the prompt vector.

2. The method for constructing a language large model for the forestry vertical field according to claim 1, wherein: In S1, the business is forest right acceptance, review, and certification, the text is the text data corresponding to various types of business, and the image is the remote sensing image and satellite image corresponding to the text.

3. A method for constructing a language large model for the forestry vertical field according to claim 1, characterized in that: In S21, the specific steps to construct an improved ViLBERT model are; Obtain a ViLBERT model including an image stream and a text stream; The image stream includes an image coding layer, N Co-TRM layers, and a first TRM layer that are sequentially arranged. The nth Co-TRM layer in the image stream is ; The text stream includes a text encoding layer, a second TRM layer, N Co-TRM layers, and a third TRM layer that are sequentially set. The nth Co-TRM layer in the text stream is ; It is a Transformer layer, including a Q matrix Q1, a K matrix, a V matrix, a multi-head attention layer, a multi-layer perceptron, and a batch normalization layer. The Q1 is from the image encoding layer, the K matrix and the V matrix are from the second TRM layer, and according to the formula obtain , and send it into the multi-head attention layer instead of Q1; It is a Transformer layer, including a Q matrix Q2, a K matrix, and a V matrix. The Q matrix is from the second TRM layer, and the K matrix and the V matrix are from the image encoding layer. And according to the formula obtain , and replace Q2 to be fed into the multi-head attention layer.

4. A method for constructing a language large model for the forestry vertical field according to claim 1, characterized in that: When training the improved ViLBERT model in S22, minimize the total loss L total Adjust the model parameters, where; , , , , , , L total In it, L MLM , , L matching are the loss functions of the MLM task, MRM task, and multi-modal alignment task respectively, and α1, α2, and α3 are the weights of L MLM , , L matching respectively; L MLM Among them, M T is the set of positions of the masked words in the text, t i is the true label of the i-th masked word, T i is the input sequence obtained by masking the text, P(t i │T i ) is the probability that the i-th masked word is predicted as t i ; In , are, respectively, the input image and the output image for improving the ViLBERT model when performing the MRM task. The input image is the image of a randomly occluded partial region in a text-image pair, and the output image is the image predicted for the occluded region based on the occluded region and the text in the text-image pair. is the L2 norm; L matching In , sigmoid function, v T , v I are the text encoding and image encoding obtained from an improved ViLBERT model for a text-image pair in D1, respectively; The similarity function for calculating similarity, w(v T , v I ) is the dynamic weighting factor, is the exp function, and β is the adjustment factor.

5. A method for constructing a language large model for the forestry vertical field according to claim 1, characterized in that: When training a large language model in S3, text data is generated using an autoregressive method, and the loss function is L generation ; , In the formula, where: t z is the z-th target word in the target text during the generation process, Z is the total number of target words in the target text, is based on the generated text and v q to calculate and generate t z the probability, and L generation The large language model is trained by maximizing the conditional probability of the target text.

6. The method for constructing a language large model for the forestry vertical field according to claim 1, wherein: S54 is specifically; If the query data is a query text, the prompt vector is , if the query data is a query image, the prompt vector is , taking T RAG as the K matrix and V matrix in the LLM decoder, the prompt vector as the Q matrix in the LLM decoder, and generating the target text by the LLM decoder; If the query data is a query image and a query text, first set the prompt vector to , and generate a text Text1 with T RAG ; concatenate Text1 with the query text, and then obtain a text encoding through the text stream of the multimodal encoding model , set the prompt vector to , and generate the target text with T RAG .

Citation Information

Patent Citations

  • Visual language navigation method for improving VLN-BERT based on enhanced end point alignment

    CN118820785A

  • Forestry resource change monitoring method based on multi-modal image generation

    CN118887550A