Language large model construction method oriented to vertical field of forestry
By improving the ViLBERT model and combining multimodal joint coding technology, the problem of insufficient image and text matching accuracy in forestry management is solved, and high-precision information understanding and the improvement of intelligent question-and-answer system are achieved.
Patent Information
- Application Number
- CN202510546350.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
In the prior art, the image and text matching accuracy is insufficient in the fields of forestry management and resource approval, resulting in inaccurate Q&A generated in smart Q&A.
A language model construction method for forestry vertical field is adopted. By improving the ViLBERT model, multimodal joint encoding is introduced, and 4 learnable weights are introduced in the Co-TRM layer. The model is trained by combining MLM, MRM and multimodal alignment task to ensure high-precision matching between images and text.
It significantly improves the accuracy of information understanding and the matching accuracy between images and text, provides stronger data support for forestry intelligent question-and-answer system, and promotes the development of forestry intelligent systems.
Smart Images

Figure CN120068947A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer models, and particularly to a method for constructing a large language model for the forestry vertical domain. Background Art
[0002] In forestry management, especially in the business process of "forest right acceptance, review, and certification", the accurate acquisition of data and information is crucial. With the rapid development of image-text fusion technology, large language models based on multimodal learning have shown great potential in the fields of forestry intelligent question answering, business process handling, and document generation. As a multimodal deep learning method, ViLBERT not only significantly improves the accuracy of information understanding by introducing joint encoding of images and text, but also ensures the matching accuracy of images and text, thus promoting the development of forestry intelligent systems.
[0003] In forestry business processing and intelligent question answering, ensuring the efficient fusion of image and text information is crucial for performing accurate question answering and document generation. Related application requirements often require the close combination of images and text to meet the needs of business process and knowledge question answering. However, the limitations of existing technologies result in insufficient matching accuracy between images and text, leading to inaccurate question answering in intelligent question answering. In forestry business management, image data (such as satellite images of land or forest areas) and text data (such as forest right documents) are closely related. How to use this feature to achieve high-precision matching of images and text and apply it to intelligent question answering is a problem that needs to be solved.
[0004] ViLBERT (Vision and Language BERT) is a multi-modal pre-training model designed to process image and text data simultaneously. It consists of two parallel BERT (Bidirectional Encoder Representation from Transformer) streams, an image stream and a text stream, for processing image data and text data respectively. Each stream is composed of multiple TRM (Transformer) and Co-TRM layers (co-attentional TRM layers). The biggest improvement of Co-TRM lies in swapping the K matrix and V matrix of the two streams, thus achieving information fusion between the two different streams. Therefore, the number of Co-TRM layers in the image stream and the text stream should be the same, and they perform information fusion in one-to-one correspondence. For example, if there are N Co-TRM layers in each stream, the nth Co-TRM layer in the image stream should correspond one-to-one with the nth Co-TRM layer in the text stream. In the corresponding two Co-TRM layers, the Q matrix of the Co-TRM layer in the image stream comes from image information, while the K matrix and V matrix come from text information. For the Co-TRM layer in the text stream, the Q matrix comes from text information, and the K matrix and V matrix come from image information.
[0005] ViLBERT uses two pre-training tasks: Task 1: Masked multi-modal modelling task, which includes MLM and MRM. MLM (Masked Language Model) aims to predict some randomly masked words in the input sentence. MRM (Masked Region Model) randomly masks some regions in the image and uses the remaining part of the image to predict the content of the masked regions.
[0006] Task 2: Multi-modal alignment prediction task. The model presents image-text pairs and predicts whether the image and text are aligned.
[0007] A large language model (LLM) refers to a deep learning model trained using a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text. Its principle is as follows: First, the input layer receives the text data input by the user and converts it into a vector representation form that the model can process, that is, performs operations such as word embedding to map words or sentences into a low-dimensional vector space. Then, the encoder based on the Transformer architecture extracts the semantic information and context information in the text. The decoder, according to the features extracted by the encoder and the previously generated text information, gradually generates the next word or character as output. By predicting the next possible vocabulary or token, the complete text content is continuously constructed to achieve tasks such as language generation and question answering. Summary of the Invention
[0008] The object of the present invention is to provide a construction method of a language large model for the forestry vertical field, which solves the problem of insufficient accuracy of existing image and text matching in the fields of forestry management and resource approval, resulting in inaccurate questions and answers generated in intelligent question answering.
[0009] In order to achieve the above object, the technical solution adopted by the present invention is as follows: A construction method of a language large model for the forestry vertical field includes the following steps; S1, construct the first dataset D1 for forestry management business Q&A; Determine multiple types of forestry management businesses. For each type of business, obtain multiple texts and the corresponding images for each text, and form text-image pairs with the corresponding texts and images. All text-image pairs form the first dataset D1; S2, construct a multi-modal encoding model, including S21~S22; S21, construct an improved ViLBERT model; Obtain a ViLBERT model, whose image stream and text stream each include N Co-TRM layers. For the nth Co-TRM layer in the image stream and the nth Co-TRM layer in the text stream , construct the first improved Q matrix according to the following formula Replace the Q matrix Q in 1 with the second improved Q matrix Replace the Q matrix Q in 2 , to obtain the improved ViLBERT model; , , In the formula, , are respectively Q in 1 and Q 2 are learnable weights, and are respectively Q in 2 and Q 1 are learnable weights, 1 ≤ n ≤ N; S22. Use D1 to train the improved ViLBERT model based on the MLM task, MRM task, and multi-modal alignment task until convergence to obtain a multi-modal encoding model. The multi-modal model is used to input image-text pairs. In the image-text pair, the image and text are respectively output as image encoding and text encoding through the image stream and text stream; S3. Train a large language model; Freeze the text stream of the multi-modal encoding model. Combine the multi-modal encoding model and the LLM decoder to form a large language model, and train the large language model to generate a text vector v q based on the user input and send it to the LLM decoder to generate the target text related to the user input when generating business intelligent answers; If the user input is text, use the multi-modal model to obtain the corresponding text encoding as v q . If the user input is an image, the multi-modal model performs a multi-modal alignment task and outputs the text encoding of the text aligned with the image as v q ; S4. Construct a RAG knowledge base D2. The k-th data in D2 is , where T k and I k are respectively the text and image in the k-th image-text pair in D1, and are respectively the text encoding and image encoding obtained by T k and I k through the multi-modal encoding model; S5. Perform intelligent answering based on RAG retrieval; S51. Input the user's query data into the multi-modal encoding model to obtain the corresponding encoding. The query data includes a query image and / or query text, and the corresponding encoding is the query image encoding and / or query text encoding ; S52. Calculate the score and weight of each data in D2. The score k and weight of the k-th data d are obtained according to the following formula; , , In the formula, γ is the score weight, is the calculation of the modulus length, is the exp function; S53, a preset threshold, filters out the data with scores greater than the threshold and places them in the candidate set B, and the data in the candidate set is weighted according to the following formula to obtain a weighted result T RAG ; , S54, generates a prompt vector based on the query data, and based on T RAG and the prompt vector, the target text is generated by the LLM decoder.
[0010] Preferably: in S1, the service is forest right acceptance, review and certification, the text is the text data corresponding to various services, and the image is the remote sensing image and satellite image corresponding to the text.
[0011] Preferably: in S21, constructing an improved ViLBERT model is specifically; Obtain a ViLBERT model including an image stream and a text stream; The image stream includes an image encoding layer, N Co-TRM layers, and a first TRM layer arranged in sequence. The nth Co-TRM layer in the image stream is ; The text stream includes a text encoding layer, a second TRM layer, N Co-TRM layers, and a third TRM layer arranged in sequence. The nth Co-TRM layer in the text stream is ; is a Transformer layer, including a Q matrix Q 1 , a K matrix, a V matrix, a multi-head attention layer, a multi-layer perceptron, and a batch normalization layer. The Q 1 is from the image encoding layer, the K matrix and the V matrix are from the second TRM layer, and according to the formula obtain , and replace Q 1 and send it into the multi-head attention layer; is a Transformer layer, including a Q matrix Q 2 , a K matrix, a V matrix, the Q matrix is from the second TRM layer, the K matrix and the V matrix are from the image encoding layer, and according to the formula obtain , and replace Q 2 and send it into the multi-head attention layer.
[0012] Preferably: when training the improved ViLBERT model in S22, the model parameters are adjusted by minimizing the total loss L total , where; , , , , , , L total In L MLM 、 、L matching are the loss functions for the MLM task, MRM task, and multi-modal alignment task respectively, and α1, α2, and α3 are the weights of L MLM 、 、L matching respectively; In L MLM , M T is the set of positions of the masked words in the text, t i is the true label of the i-th masked word, T i is the input sequence obtained by masking the text, and P(t i │T i ) is the probability that the i-th masked word is predicted as t i ; In 、 are the input image and output image for improving the ViLBERT model when performing the MRM task respectively. The input image is an image with some randomly occluded regions in a text-image pair, and the output image is an image predicted based on the occluded regions and the text in the text-image pair, is the L2 norm; In L matching : is the sigmoid function, v T 、v I are the text encoding and image encoding obtained by applying the improved ViLBERT model to a text-image pair in D1 respectively; is the similarity function for calculating similarity, w(v T , v I ) is the dynamic weighting factor, is the exp function, and β is the adjustment factor.
[0013] Preferably: When training the large language model in S3, the text data is generated by the autoregressive method, and the loss function is L generation ; , In the formula, where: t zis the z-th target word in the target text during the generation process, where Z is the total number of target words in the target text, is based on the generated text and v q calculate to generate t z probability, and L generation Train the large language model by maximizing the conditional probability of the target text.
[0014] Preferably: S54 specifically is: If the query data is a query text, the prompt vector is , if the query data is a query image, then the prompt vector is , take T RAG as the K matrix and V matrix in the LLM decoder, the prompt vector as the Q matrix in the LLM decoder, and generate the target text by the LLM decoder; If the query data is a query image and a query text, first set the prompt vector to , and generate a text Text with T RAG ; Concatenate Text 1 with the query text, and then obtain a text encoding through the text stream of the multimodal model 1 , set the prompt vector to , and generate the target text with T RAG matching .
[0015] Compared with the prior art, the advantages of the present invention are as follows: (1) According to the characteristic that in forestry management, image data (such as satellite images of land or forest areas) is closely related to text data (such as certain types of documents), introduce multimodal joint encoding, improve the ViLBERT model, and introduce 4 learnable weights in the Co-TRM layer, which can not only significantly improve the accuracy of information understanding, but also ensure the matching accuracy of images and texts, provide stronger data support for the forestry intelligent question-answering system, and thus promote the development of the forestry intelligent system.
[0016] (2) When training the improved ViLBERT model, introduce a dynamic weighting factor w(v matching ),v T ,v I ) into the loss function L T ,v I ) of its multimodal alignment task. When the similarity is high, w(v T ,v I ) will approach 1, indicating that the text and image have a high matching degree, and the model should quickly update the parameters. When the similarity is low, w(v T ,v I ) will approach 0, indicating that the matching degree of the text and image is low, and the model should slow down the update speed. In this way, the text-image pairs with a larger similarity are given higher weights, thus accelerating the training.
[0017] (3) By freezing the text stream part of the multi-modal encoding model and training a large language model based on the output of the text stream for text generation, the consistency and business relevance of text generation are ensured. Especially in the generation of forestry documents, this method provides an innovative technical path for high-quality text output that meets actual needs.
[0018] (4) In the RAG retrieval of the present invention, multiple results are weighted and fused, and then a prompt vector is generated according to the query data, and target text generation is performed together with the weighted results of the RAG retrieval. Since a prompt vector can be obtained regardless of whether the input query data is a query text, a query image, or a combination of both, the flexibility of retrieval is improved, and while ensuring flexibility, the accuracy of the target text is also taken into account.
[0019] In summary, the present invention has greatly improved the quality and application value of forestry intelligent question answering and document generation by introducing innovative technologies such as multi-modal joint encoding, weighted similarity calculation, and freezing the text encoder to train the LLM, providing an innovative solution for forestry-related fields. By solving the problem of insufficient matching accuracy between existing images and texts, this method provides a new technical path for forestry business processes, intelligent question answering, and document generation, and has broad application prospects especially in various forestry management business types. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a structural diagram of the improved ViLBERT model of the present invention; Figure 2 In the improved ViLBERT model and Interaction schematic diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The present invention will be further described below in conjunction with embodiments and drawings.
[0022] Embodiment 1: Refer to Figure 1 and Figure 2 A method for constructing a language large model for the forestry vertical field includes the following steps; S1, constructing a first data set D1 for forestry management business question answering; Determine multiple types of forestry management businesses. For each type of business, obtain multiple texts and the corresponding images for each text, form text-image pairs with the corresponding texts and images, and all text-image pairs form the first data set D1; S2, constructing a multi-modal encoding model, including S21~S22; S21, constructing an improved ViLBERT model; Obtain a ViLBERT model, where both the image stream and the text stream include N Co-TRM layers. For the nth Co-TRM layer in the image stream and the nth Co-TRM layer in the text stream , construct the first improved Q matrix according to the following formula Replace the Q matrix Q in 1 , the second improved Q matrix Replace the Q matrix Q in 2 to obtain an improved ViLBERT model; , , In the formula, , are respectively the learnable weights of Q in 1 , Q 2 , 1 ≤ n ≤ N; , are respectively the learnable weights of Q in 2 , Q 1 , 1 ≤ n ≤ N; S22, use D1 to train the improved ViLBERT model based on the MLM task, MRM task, and multi-modal alignment task until convergence to obtain a multi-modal encoding model. The multi-modal model is used to input image-text pairs, and the image and text in the image-text pair are respectively output as image encoding and text encoding through the image stream and the text stream; S3, train a large language model; Freeze the text stream of the multi-modal encoding model, and form a large language model by combining the multi-modal encoding model and the LLM decoder, and train the large language model to generate a text vector v q according to the user input and send it to the LLM decoder to generate the target text related to the user input when generating business intelligent answers; If the user input is text, obtain the corresponding text encoding as v q by the multi-modal model. If the user input is an image, the multi-modal model performs a multi-modal alignment task and outputs the text encoding of the text aligned with the image as v q ; S4, construct a RAG knowledge base D2, and the kth data in D2 is , where T k , I k are respectively the text and image in the kth image-text pair in D1, , are respectively T k , I kThe text encoding and image encoding obtained by the multi-modal encoding model; S5. Perform intelligent question answering based on RAG retrieval; S51. Input the user's query data into the multi-modal encoding model to obtain corresponding encodings. The query data includes query images and / or query texts, and the corresponding encodings are query image encodings and / or query text encodings ; S52. Calculate the scores and weights of each piece of data in D2. The score k and weight of the k-th piece of data d are obtained according to the following formula; , , In the formula, γ is the score weight, is to calculate the norm, is the exp function; S53. Preset a threshold, filter out the data with scores greater than the threshold and place them in the candidate set B, and obtain a weighted result T by weighting the data in the candidate set according to the following formula RAG ; , S54. Generate a prompt vector according to the query data, and generate the target text by the LLM decoder based on T RAG and the prompt vector.
[0023] In this embodiment S1, the service is forest right acceptance, review and certification. The text is the text data corresponding to various services, and the image is the remote sensing image and satellite image corresponding to the text.
[0024] In S21, specifically constructing an improved ViLBERT model is as follows; Obtain a ViLBERT model including an image stream and a text stream; The image stream includes an image encoding layer, N Co-TRM layers, and a first TRM layer arranged in sequence. The n-th Co-TRM layer in the image stream is ; The text stream includes a text encoding layer, a second TRM layer, N Co-TRM layers, and a third TRM layer arranged in sequence. The n-th Co-TRM layer in the text stream is ; is a Transformer layer, including a Q matrix Q 1 , a K matrix, a V matrix, a multi-head attention layer, a multi-layer perceptron, and a batch normalization layer. The Q 1 comes from the image encoding layer, the K matrix and the V matrix come from the second TRM layer, and according to the formula Obtain , replacing Q 1 and sending it into the multi-head attention layer; is a Transformer layer, including Q matrix Q 2 , K matrix, V matrix. The Q matrix is from the second TRM layer, and the K matrix and V matrix are from the image encoding layer. According to the formula obtain , replacing Q 2 and sending it into the multi-head attention layer. In Figure 1 , use to replace Q 1 After that, the Co-TRM layer in the original image stream is called the first improved Co-TRM layer and uses to replace Q 2 After that, the Co-TRM layer in the original text stream is called the second improved Co-TRM layer.
[0025] When training the improved ViLBERT model in S22, adjust the model parameters to minimize the total loss L total , where; , , , , , , L total In L MLM , , L matching are the loss functions of the MLM task, MRM task, and multi-modal alignment task respectively. α1, α2, α3 are the weights of L MLM , , L matching respectively; L MLM In M T is the set of positions of the masked vocabulary in the text, t i is the true label of the i-th masked vocabulary, T i is the input sequence obtained by masking the text, P(t i │T i ) is the probability that the i-th masked vocabulary is predicted as t i ; In , They are the input image and the output image for improving the ViLBERT model when performing the MRM task. The input image is the image of a randomly occluded partial area in a text-image pair, and the output image is the image predicted for the occluded area based on the occluded area and the text in the text-image pair. is the L2 norm; L matching in is the sigmoid function, and v T , v I are respectively the text encoding and the image encoding obtained by an improved ViLBERT model for a text-image pair in D1; is the similarity function for calculating similarity, w(v T , v I ) is the dynamic weighting factor, is the exp function, and β is the adjustment factor.
[0026] When training the large language model in S3, text data is generated using the autoregressive method, and the loss function is L generation ; , In the formula, where: t z is the z-th target word in the target text during the generation process, and Z is the total number of target words in the target text, is the probability of generating t q based on the generated text and v z , and L generation trains the large language model by maximizing the conditional probability of the target text.
[0027] Specifically, S54 is as follows: If the query data is a query text, the prompt vector is , if the query data is a query image, the prompt vector is , take T RAG as the K matrix and the V matrix in the LLM decoder, the prompt vector as the Q matrix in the LLM decoder, and generate the target text by the LLM decoder; if the query data is a query image and a query text, first set the prompt vector to , generate a text Text RAG with T 1 ; Concatenate Text 1 with the query text, and then obtain a text encoding through the text stream of the multimodal model, set the prompt vector to , and generate the target text with T RAG .
[0028] Regarding the training of the multimodal encoding model, through the following three tasks: MLM Task: In an intelligent question-answering system related to forestry, it is necessary to transform the user's query into relevant business content, and the professionalism of forestry knowledge should be taken into account during this process. During the training process, through the masking mechanism, real text and image missing scenarios can be simulated, thereby enhancing the robustness of the model. By masking the key information in the business-related text, the model can learn to correctly fill in the missing words given the context.
[0029] MRM Task: The masking loss of the image part is very important for training the model to understand image details. For example, in forest resource management, some areas of satellite images may be lost or occluded, and the model needs to be able to infer the content of the missing part from the known areas.
[0030] Multimodal Alignment Task: In the "intelligent question-answering" session, the image and text are matched to ensure the semantic consistency of the input image and the corresponding text. Through dynamic weighting, the image-text pairs with higher similarity will be assigned higher weights, thus accelerating the training.
[0031] Regarding the improved large language model: The large language model is an encoder-decoder architecture. Before training the large language model, this invention first uses a multimodal encoding model as the new encoder to replace the original encoder of the large language model, and forms an improved large language model by combining the multimodal encoding model and the LLM decoder, and then trains it according to the existing technical methods. Freezing the text stream part of the multimodal encoding model means that the parameters of the text stream will not be updated during the subsequent training process. When the multimodal encoding model inputs an image or text, it will generate a text vector according to step S3 , and the LLM decoder will generate the target text related to the query data according to .
[0032] Regarding RAG Retrieval: In various forestry management-related intelligent question-answering services, it is necessary to efficiently retrieve a large amount of image-text data to quickly provide accurate query results for users. This invention weights and fuses multiple data obtained by RAG retrieval to generate a unified semantic representation, which is used as the input of the text generation model together with the prompt vector for the generation of the target text, and has the characteristics of flexible modal input, unified and accurate retrieval expression. Based on the RAG-based retrieval mechanism, this invention can perform intelligent screening according to the similarity between text and image, so as to provide accurate responses for the "intelligent question-answering" system.
[0033] Regarding the LLM decoder and its applications: In large language models, especially in models based on the Transformer architecture, the Q matrix, K matrix, and V matrix are core components in the self-attention mechanism. When training the LLM decoder in S3, the Q matrix, K matrix, and V matrix are directly obtained from text vectors and the target text is finally generated. When applying the LLM decoder subsequently, in step S5, the prompt vector and T are combined RAG Generate the Q matrix, K matrix, and V matrix to obtain the target text.
[0034] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing a large language model for the forestry vertical field, characterized by: The steps include: S1, construct the first data set D1 of forestry management business questions and answers; Determine multiple types of forestry management services, obtain multiple texts and images corresponding to each text for each type of service, form image-text pairs with the corresponding texts and images, and all image-text pairs form a first data set D1; S2, construct a multimodal coding model, including S21~S22; S21, construct an improved ViLBERT model; Get a ViLBERT model, whose image stream and text stream each include N Co-TRM layers, and for the nth Co-TRM layer in the image stream and the nth Co-TRM layer in the text stream , the first improved Q matrix is constructed according to the following formula replace Middle Q matrix Q1, second improved Q matrix replace The Q matrix Q2 in the middle obtains the improved ViLBERT model; , , In the formula, , They are The learnable weights of Q1 and Q2, , They are The learnable weights of Q2 and Q1 are 1≤n≤N; S22, using D1 to train the improved ViLBERT model based on the MLM task, the MRM task, and the multimodal alignment task until convergence, to obtain a multimodal encoding model, wherein the multimodal model is used to input an image-text pair, wherein the image and the text in the image-text pair are outputted via an image stream and a text stream, respectively, to output an image code and a text code; S3, train a large language model; Freeze the text stream of the multimodal encoding model, form a large language model with the multimodal encoding model and the LLM decoder, and train the large language model to generate text vectors v according to user input q , sent to the LLM decoder to generate the target text related to the user input when answering business intelligent questions; If the user input is text, the corresponding text encoding is obtained by the multimodal model as v q , if the user input is an image, the multimodal model performs the multimodal alignment task and outputs the text encoding of the text aligned with the image as v q ; S4, construct the RAG knowledge base D2, the kth data in D2 is , where T k ,I k are the text and image in the kth image-text pair in D1, , T k ,I k Text encoding and image encoding obtained by multimodal encoding model; S5, intelligent question answering based on RAG retrieval; S51, inputting the user's query data into the multimodal coding model to obtain a corresponding code, wherein the query data includes a query image and / or a query text, and the corresponding code is a query image code and / or query text encoding ; S52, calculate the score and weight of each data in D2, where the kth data d k Score and weight According to the following formula: , , In the formula, γ is the score weight, To calculate the modulus length, is the exp function; S53, preset a threshold, filter out data with scores greater than the threshold and place them in candidate set B, and weight the data in the candidate set according to the following formula to obtain a weighted result T RAG ; , S54, generating a prompt vector according to the query data, based on T RAG and the hint vector, the target text is generated by the LLM decoder.
2. The method for constructing a large language model for the forestry vertical field according to claim 1 is characterized by: In S1, the business is forest rights acceptance, review and certification, the text is text data corresponding to various businesses, and the image is the remote sensing image and satellite image corresponding to the text.
3. The method for constructing a large language model for the forestry vertical field according to claim 1 is characterized in that: In S21, an improved ViLBERT model is constructed as follows; Obtain a ViLBERT model including an image stream and a text stream; The image stream includes an image coding layer, N Co-TRM layers, and a first TRM layer arranged in sequence. The nth Co-TRM layer in the image stream is ; The text stream includes a text coding layer, a second TRM layer, N Co-TRM layers, and a third TRM layer, which are arranged in sequence. The nth Co-TRM layer in the text stream is ; is a Transformer layer, including Q matrix Q1, K matrix, V matrix, multi-head attention layer, multi-layer perceptron and batch normalization layer. The Q1 comes from the image coding layer, the K matrix and V matrix come from the second TRM layer, and according to the formula get , replace Q1 and send it to the multi-head attention layer; is a Transformer layer, including Q matrix Q2, K matrix, V matrix, Q matrix comes from the second TRM layer, K matrix and V matrix come from the image coding layer, and according to the formula get , replacing Q2 and sending it to the multi-head attention layer.
4. The method for constructing a large language model for the forestry vertical field according to claim 1 is characterized by: When training the improved ViLBERT model in S22, the total loss L is minimized total Adjust model parameters, where; , , , , , , L total In, L MLM , , L matching They are the loss functions of the MLM task, MRM task, and multimodal alignment task, respectively. α1, α2, and α3 are L MLM , , L matching The weight of L MLM In, M T is the set of positions of the masked words in the text, t i is the true label of the i-th masked word, T i is the input sequence obtained by masking the text, P(t i │T i ) is the i-th masked word predicted as t i The probability of middle, , are respectively the input image and output image of the improved ViLBERT model when performing the MRM task, wherein the input image is an image of a randomly occluded area in an image-text pair, and the output image is an image predicted based on the occluded area and the text in the image-text pair. is the L2 norm; L matching middle, is the sigmoid function, v T 、v I They are the text encoding and image encoding of a picture-text pair in D1 obtained by the improved ViLBERT model; is the similarity function for calculating similarity, w(v T ,v I ) is the dynamic weighting factor, is the exp function, and β is the adjustment factor.
5. The method for constructing a large language model for the forestry vertical field according to claim 1 is characterized in that: When training a large language model in S3, the autoregressive method is used to generate text data, and the loss function is L generation ; , In the formula, where: t z is the zth target word in the target text during the generation process, Z is the total number of target words in the target text, Based on the generated text and v q Calculate the generated t z The probability that L generation Large language models are trained by maximizing the conditional probability of the target text.
6. The method for constructing a large language model for the forestry vertical field according to claim 1 is characterized by: S54 is as follows: If the query data is query text, the hint vector is , if the query data is a query image, then the hint vector is , T RAG As the K matrix and V matrix in the LLM decoder, the prompt vector is used as the Q matrix in the LLM decoder, and the LLM decoder generates the target text; If the query data is the query image and query text, first let the prompt vector be , and T RAG Generate a text Text1; concatenate Text1 with the query text, and then obtain a text encoding through the multimodal model text stream , let the prompt vector be , and T RAG Generate target text.
Citation Information
Patent Citations
Content searching method and related device
CN117891980A
Visual language navigation method for improving VLN-BERT based on enhanced end point alignment
CN118820785A
Forestry resource change monitoring method based on multi-modal image generation
CN118887550A
Method and system for visio-linguistic understanding using contextual language model reasoners
US20220019734A1
Cited By
Intelligent customer service dialogue generation method and system based on large model
CN120873126A
Forest management data processing method driven by large language model
CN120994908A
Scientific research reasoning task processing method, equipment and medium
CN121480697A