Forestry image generation method with visual prompt

By using a large language model with a knowledge enhancement module based on forestry concept system in forestry image generation, standard instance descriptions conforming to forestry survey specifications are generated and ranked according to intent relevance. This solves the problem of visual cues deviating from user intent in existing technologies and achieves more accurate forestry image generation.

CN120894682APending Publication Date: 2025-11-04NANJING FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510775933.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing methods for generating forestry images with visual cues suffer from problems such as deviating from the user's intent and insufficiently accurate cues. In particular, when the user's input of natural language is ambiguous, it is difficult to accurately control the retrieval and recognition range.

Method used

The large language model, which is pre-linked with knowledge enhancement modules of forestry concept system, is used to recognize the input language, generate standard instance descriptions that conform to forestry survey specifications, prioritize them according to the degree of intent relevance, and divide image regions for visual cues labeling.

Benefits of technology

It improves the accuracy and consistency of visual cues with intent, and the generated forestry images are closer to the user's needs, balancing precision and intent accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a forestry image generation method with a visual prompt in the technical field of image recognition, and aims to solve the technical problems that the intention of a user is deviated and the prompt is not accurate enough. The method comprises the following steps: firstly, acquiring contents input by a user through voice or characters, performing recognition by utilizing a large language model pre-associated with a forestry concept system knowledge enhancement module, and converting a non-standard natural language into standard instance description conforming to forestry survey specifications; part of the descriptions accord with the real intention of the user, and part of the descriptions have large differences. The standard instance description is subjected to priority ranking according to the intention association degree with the input language, so that the description closer to the intention of the user is screened, and the accuracy of visual prompt marking is improved. Image areas are divided from the forestry image according to each standard instance description, visual prompt labeling is carried out on the areas according to priorities, and finally the forestry image with visual prompts is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a forestry image generation method with visual prompts, and belongs to the technical field of image recognition. BACKGROUND

[0002] In the field of forestry investigation, visual recognition models have also been used to identify and label information on forestry images.

[0003] When using a visual recognition model to identify a forestry image, the natural language input by the user may ignore many features that should be expressed in the description, resulting in a gap between the visual prompts given by the visual recognition model and the user's intention when identifying the forestry image. When the natural language input by the user is ambiguous, it is difficult for the visual recognition model to accurately control the search and recognition range when identifying the forestry image, resulting in an excessively large search and recognition range, which labels both high and low correlation content on the forestry image.

[0004] Therefore, the existing forestry image generation method with visual prompts has the problems of deviating from the user's intention and not being accurate enough. SUMMARY

[0005] The present application aims to overcome the deficiencies in the prior art and provide a forestry image generation method with visual prompts that is closer to the user's intention and more accurate.

[0006] To achieve the above-mentioned purpose, the present application is implemented by using the following technical solutions:

[0007] In a first aspect, the present application provides a forestry image generation method with visual prompts, comprising,

[0008] obtaining an input language;

[0009] identifying the input language through a large language model pre-associated with a forestry concept system knowledge enhancement module to obtain a plurality of standard instance descriptions conforming to forestry investigation specifications;

[0010] prioritizing all standard instance descriptions according to their relevance to the intention of the input language;

[0011] dividing image regions from the forestry image according to each standard instance description;

[0012] labeling the image regions with visual prompts according to the priority of each standard instance description to obtain a forestry image with visual prompts.

[0013] In some embodiments of the first aspect, the forestry image is pre-divided into sub-regions based on forest region categories to obtain sub-region images, each sub-region image is labeled with a corresponding forest region category;

[0014] The visual cue annotation is performed on each partition image respectively, and the partition images with visual cue annotation are fused to obtain a forestry image with visual cue;

[0015] The priority of all standard instance descriptions is ranked according to the intention association degree of the input language, including,

[0016] The priority of all standard instance descriptions is ranked according to the intention association degree of the input language and the association degree of the language concept under the forest area category.

[0017] In some embodiments of the first aspect, the input language is identified by the large language model pre-associated with the forestry concept system knowledge enhancement module, and a plurality of standard instance descriptions conforming to the forestry survey specification are obtained, including,

[0018] The large language model extracts a nominal phrase from the input language, and arranges the nominal phrase into a candidate instance set C={c1, c2,..., cn}, and the candidate instance set is obtained as follows:

[0019] ,

[0020] In the formula, is a large language model, represents an extraction action, is an input language, c is a candidate instance, n is the number of candidate instance sets, C is a candidate instance set, c1 and c2 represent the first and second candidate instances in the candidate instance set;

[0021] The large language model uses the forestry concept system knowledge enhancement module to perform semantic expansion and context association on the candidate instance, generates a plurality of standard instance descriptions, and arranges the standard instance descriptions into a first instance set The first instance set is used to obtain a forestry image with visual cue, and the first instance set is obtained as follows:

[0022] ;

[0023] In the formula, is a first instance set, m is the number of standard instance descriptions, and represent the first and second standard instance descriptions, represents a generation action.

[0024] In some embodiments of the first aspect, the input language is identified by the large language model pre-associated with the forestry concept system knowledge enhancement module, and a plurality of standard instance descriptions conforming to the forestry survey specification are obtained, further including,

[0025] The large language model obtains the query intent based on the input language, and filters the first instance set according to the query intent to obtain the second instance set. The second set of instances is used to obtain forestry images with visual cues;

[0026] In the formula, For the second set of instances, The number of standard instance descriptions in the second instance set.

[0027] In some embodiments of the first aspect, the step of dividing the image region from the forestry image according to each standard instance description includes,

[0028] Using image processing tools, the forestry image is divided into multiple image regions based on the second instance set. Each standard instance description corresponds to one or more image regions;

[0029] Image regions corresponding to the same standard instance are integrated into a subset of regions:

[0030] ,

[0031] In the formula, The serial number described for the standard instance. For the corresponding number A subset of regions described by a standard instance, For image region, This is the index of the image region within the subset. This represents the number of image regions within the subset. For the corresponding number The first, second... of the regions described in the standard instance subset Image regions;

[0032] Integrate the subsets of regions corresponding to different standard instances into a region set:

[0033] ,

[0034] In the formula, For a set of regions, The number of region subsets, For the 1st, 2nd... in the set of regions A subset of regions;

[0035] Construct the mapping relationship between the second set of instances and the set of regions:

[0036] ;

[0037] The image region is visually prompted and labeled according to the priority of each standard instance description to obtain a forestry image with visual prompts, including,

[0038] An image processing tool is obtained, which generates a bounding box covering the image region on the forestry image by the following formula:

[0039] ,

[0040] In the formula, is the forestry image with a bounding box, is the input image, is the operation of the image processing tool.

[0041] In some embodiments of the first aspect, the image processing tool obtains the forestry image with a bounding box by the following formula, including,

[0042] All bounding boxes are sorted according to the pixel area;

[0043] The bounding boxes are generated in order from large to small according to the pixel area;

[0044] In response to the bounding box with relatively late generation order overlapping with the bounding box with relatively early generation order, the position of the bounding box with relatively late generation order is adjusted until no longer overlapping;

[0045] The bounding boxes corresponding to different standard instance descriptions are displayed in different colors.

[0046] In some embodiments of the first aspect, the forestry image with visual prompts is obtained by visually prompting and labeling the image region according to the priority of each standard instance description, further including,

[0047] Using the large language model to generate label information corresponding to all standard instance descriptions;

[0048] The label information is sorted into a label set:

[0049] , wherein the label corresponding to the i-th image region in the i-th standard instance description description corresponding region subset,

[0050] In the formula, is the label set, is the i-th label subset in the label set, is the i-th label subset in the label set;

[0051] ​​​​Construct a mapping relationship between the label set and the region set:

[0052] ;

[0053] Take the label as input text , use a visual language model to obtain output text :

[0054] ,

[0055] In the formula, is the input text, is the forestry image, is the output text;

[0056] Display the output text on the bounding box as a forestry image with visual cues.

[0057] In a second aspect, the present application also provides a computer device, comprising a processor and a memory connected to the processor, and the memory stores a computer program, when the computer program is executed by the processor, the steps of the forestry image generation method with visual cues according to any one of the embodiments of the first aspect are executed.

[0058] In a third aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to realize the steps of the forestry image generation method with visual cues according to any one of the embodiments of the first aspect.

[0059] In a fourth aspect, the present application also provides a computer program product, comprising computer programs / instructions, characterized in that the computer programs / instructions are executed by a processor to realize the steps of the forestry image generation method with visual cues according to any one of the embodiments of the first aspect.

[0060] Compared with the prior art, the present application has the following beneficial effects:

[0061] The forestry image with visual cues provided by the application is generated by first obtaining the content input by the user through voice or text, recognizing the non-standard natural language by using the large language model pre-associated with the forestry concept system knowledge enhancement module, and converting the non-standard natural language into standard instance descriptions conforming to the forestry survey specification. Some of these descriptions conform to the user's true intention, and some have a large gap. Then, the standard instance descriptions are prioritized according to the intention correlation with the input language to filter the descriptions closer to the user's intention and improve the accuracy of visual cue labeling. Then, according to each standard instance description, the image area is divided from the forestry image, and the visual cue labeling is performed on these areas according to the priority, and finally the forestry image with visual cues is obtained. This method takes into account the accuracy and intention accuracy, so that the generated image is more in line with the user's needs. BRIEF DESCRIPTION OF DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0063] Figure 1 is the step flow chart of the forestry image with visual cues provided by the embodiment;

[0064] Figure 2 is the step flow chart of the forestry image with visual cues from another angle;

[0065] Figure 3 is the effect schematic diagram of the bounding box generated by using the forestry image with visual cues provided by the embodiment;

[0066] Figure 4 is the effect schematic diagram of visual cue labeling on the forestry image on a sunny day by using the forestry image with visual cues provided by the embodiment;

[0067] Figure 5 is the effect schematic diagram of visual cue labeling on the forestry image on a cloudy day by using the forestry image with visual cues provided by the embodiment;

[0068] Figure 6 is the principle schematic diagram of the computer device provided by the embodiment. DETAILED DESCRIPTION

[0069] The technical solutions of the present application will be described in detail below with reference to the drawings and specific embodiments. It should be understood that the embodiments and specific features in the embodiments are detailed descriptions of the technical solutions of the present application, and are not limitations of the technical solutions of the present application. In the case of no conflict, the technical features in the embodiments and the embodiments can be combined with each other.

[0070] The term "and / or" herein is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent three cases of A alone, A and B together, and B alone. In addition, the character " / " herein generally represents that the associated objects before and after are in an "or" relationship.

[0071] Embodiment one:

[0072] Figure 1 is a flowchart of a forestry image generation method with visual prompts in the first embodiment of the present application. The flowchart only shows the logical order of the method described in the embodiment. In other possible embodiments of the present application, the steps shown or described can be completed in an order different from that shown in the embodiment on the premise that they do not conflict with each other. Figure 1

[0073] The forestry image generation method with visual prompts provided in the embodiment can be applied to a terminal, such as any smart phone, tablet computer or computer device with communication function. Referring to Figure 1 , the method of the embodiment specifically includes the following steps:

[0074] Obtaining an input language; that is, a user inputs content to execute the main body of the method through voice or text;

[0075] Recognizing the input language through a large language model pre-associated with a forestry concept system knowledge enhancement module to obtain a plurality of standard instance descriptions conforming to forestry survey specifications; the large language model pre-associated with the forestry concept system knowledge enhancement module can load the training forestry concept system knowledge enhancement module on the large language model, or a commercial large language model can be used; this step aims to expand the input language through the recognition ability of the large language model on human natural language, and also aims to convert non-standard natural language into machine language conforming to forestry survey specifications. Some of the obtained standard instance descriptions contain the real intention of the user, and others are quite different from the real intention of the user;

[0076] Prioritizing all standard instance descriptions according to the intention correlation with the input language; aiming to continue to use the advantage of the large language model in analyzing human natural language to screen the standard instance descriptions, improve the accuracy of visual prompt labeling in the later stage, and close to the user's intention; ​

[0077] According to each standard instance description, image regions are divided from the forestry image;

[0078] The image regions are visually prompted according to the priority of the standard instance description, and a forestry image with visual prompts is obtained, so that the standard instance description close to the user's intention is given priority to visual prompt, and the overall accuracy and intention accuracy of the forestry image with visual prompts are considered.

[0079] The forestry image with visual prompts provided by the embodiment is generated by first obtaining the content input by the user through voice or text, recognizing it using a large language model pre-associated with a forestry concept system knowledge enhancement module, and converting non-standard natural language into standard instance descriptions conforming to forestry survey specifications. Some of these descriptions conform to the user's true intention, and some have a large gap. Then, the standard instance descriptions are prioritized according to the intention correlation with the input language to filter descriptions closer to the user's intention and improve the accuracy of visual prompt labeling. Then, according to each standard instance description, image regions are divided from the forestry image, and the regions are visually prompted according to the priority, and finally a forestry image with visual prompts is obtained. This method takes into account accuracy and intention accuracy, making the generated image more in line with user needs.

[0080] Embodiment Two:

[0081] The embodiment provides a forestry image with visual prompts. The embodiment is optimized based on embodiment one to improve technical effects and refine technical solutions. Details not described in this embodiment are described in embodiment one.

[0082] The embodiment notes that in the same forestry aerial photograph or satellite image, there may be multiple categories such as protective forest, timber forest, economic forest, fuel forest, and special-purpose forest. Different categories correspond to different sub-branches of the forestry concept system. If the partitioned images of different categories are uniformly visually prompted, the accuracy will be reduced. Therefore, the forestry image is pre-divided into images based on the forest area category to obtain partitioned images, each of which is labeled with the corresponding forest area category; each partitioned image is visually prompted, and the large language model performs parallel operation on each partitioned image, and finally the partitioned images with visual prompts are fused to obtain a complete forestry image with visual prompts. The priority of all standard instance descriptions is sorted according to the intention correlation with the input language, including sorting the priority of all standard instance descriptions according to the intention correlation with the input language and the correlation with the language concept under the forest area category, so that the visual prompt of different forest area categories is more accurate and closer to the user's original intention.

[0083] As one of the embodiments, the large language model extracts a nominal phrase from the input language, the nominal phrase being the basis for semantic association of the large language model, and arranges the nominal phrase into a candidate instance set C = {c1, c2, …, cn}, the candidate instance set being obtained as follows:

[0084] ,

[0085] In the formula, is a large language model, represents an extraction action, is an input language, c is a candidate instance, n is the number of candidate instances, C is a candidate instance set, c1 and c2 represent the first and second candidate instances in the candidate instance set;

[0086] The large language model uses the forest concept system knowledge enhancement module to perform semantic expansion and context-related association on the candidate instances, generates multiple standard instance descriptions, and arranges the standard instance descriptions into a first instance set , the first instance set being used to obtain forestry images with visual prompts, the first instance set being obtained as follows:

[0087] ;

[0088] In the formula, is a first instance set, m is the number of standard instance descriptions, and represent the first and second standard instance descriptions, represents a generation action.

[0089] As one of the embodiments, the large language model associated with the forest concept system knowledge enhancement module identifies the input language and obtains multiple standard instance descriptions that meet the forestry survey specifications, and further includes,

[0090] The large language model obtains a query intent according to the input language, filters the first instance set according to the query intent, and obtains a second instance set , the second instance set being used to obtain forestry images with visual prompts;

[0091] In the formula, is a second instance set, is the number of standard instance descriptions in the second instance set.

[0092] The standard instance descriptions generated by the large language model are generally more than the nominal phrases in the input language. The step aims to use the understanding ability of the large language model for human language to preliminarily screen the standard instance descriptions, and eliminate the standard instance descriptions that obviously deviate from the description intention of the user. In the later priority sorting, even the standard instance descriptions with poor priority also have certain closeness to the real semantics of the user.

[0093] As one of the embodiments, the image region is divided from the forestry image according to each standard instance description, including,

[0094] The forestry image is divided into a plurality of image regions according to the second instance set by using an image processing tool Each standard instance description corresponds to one or more image regions;

[0095] The image regions corresponding to the same standard instance description are integrated into a region subset:

[0096] ,

[0097] In the formula, is the count serial number of the standard instance description, is the region subset corresponding to the first standard instance description, is the image region, is the count serial number of the image region in the region subset, is the number of image regions in the region subset, is the first, second... image region in the region subset corresponding to the first standard instance description;

[0098] The region subsets corresponding to different standard instance descriptions are integrated into a region set:

[0099] ,

[0100] In the formula, is the region set, is the number of region subsets, is the first, second... region subset in the region set;

[0101] The mapping relationship between the second instance set and the region set is constructed:

[0102] ;

[0103] The image region is visually prompted and labeled according to the priority of each standard instance description, and a forestry image with visual prompt is obtained, including,

[0104] The image processing tool generates a bounding box covering the image region on the forestry image by the following formula:

[0105] ,

[0106] In the formula, is the forestry image of the bounding box, is the input image, is the operation of the image processing tool.

[0107] The present embodiment constructs a mapping relationship between the image region and the standard instance description, so that when the image processing tool generates a bounding box (i.e. a visual prompt) covering the image region on the forestry image, it has the basis for generating the bounding box according to the priority.

[0108] In addition to generating according to priority, the generation of the image box also needs to be generated in order according to size. As one of the embodiments, the image processing tool obtains the forestry image with the bounding box by the following formula, including,

[0109] Sort all the bounding boxes according to the pixel area;

[0110] Generate the bounding boxes in order from large to small according to the pixel area; that is, in the case of the same priority of the corresponding standard instance description, the bounding box with larger pixel area is generated first;

[0111] In response to the overlapping of the bounding box with relatively late generation order and the bounding box with relatively early generation order, the position of the bounding box with relatively late generation order is adjusted until there is no longer overlapping;

[0112] The bounding boxes corresponding to different standard instance descriptions are displayed in different colors. The forestry image with visual prompts is clean and neat, and the visual prompts are labeled from heavy to light,

[0113] As one of the embodiments, the image region is visually prompted and labeled according to the priority of each standard instance description to obtain the forestry image with visual prompts, further including,

[0114] Generate the label information corresponding to all standard instance descriptions using the large language model;

[0115] Organize the label information into a label set:

[0116] , wherein the label corresponding to the image region in the th image region in the region subset corresponding to the th standard instance description,

[0117] In the formula, The tag set is constructed as a tag set, The first, second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, eleventh, twelfth, thirteenth, fourteenth, fifteenth, sixteenth, seventeenth, eighteenth, nineteenth, twentieth, twenty-first, twenty-second, twenty-third, twenty-fourth, twenty-fifth, twenty-sixth, twenty-seventh, twenty-eighth, twenty-ninth, thirtieth, thirty-first, thirty-second, thirty-third, thirty-fourth, thirty-fifth, thirty-sixth, thirty-seventh, thirty-eighth, thirty-ninth, or fortieth tag subset of the tag set, The first, second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, eleventh, twelfth, thirteenth, fourteenth, fifteenth, sixteenth, seventeenth, eighteenth, nineteenth, twentieth, twenty-first, twenty-second, twenty-third, twenty-fourth, twenty-fifth, twenty-sixth, twenty-seventh, twenty-eighth, twenty-ninth, thirtieth, thirty-first, thirty-second, thirty-third, thirty-fourth, thirty-fifth, thirty-sixth, thirty-seventh, thirty-eighth, thirty-ninth, or fortieth tag subset of the tag set, The first, second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, eleventh, twelfth, thirteenth, fourteenth, fifteenth, sixteenth, seventeenth, eighteenth, nineteenth, twentieth, twenty-first, twenty-second, twenty-third, twenty-fourth, twenty-fifth, twenty-sixth, twenty-seventh, twenty-eighth, twenty-ninth, thirtieth, thirty-first, thirty-second, thirty-third, thirty-fourth, thirty-fifth, thirty-sixth, thirty-seventh, thirty-eighth, thirty-ninth, or fortieth tag subset of the tag set; The first, second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, eleventh, twelfth, thirteenth, fourteenth, fifteenth, sixteenth, seventeenth, eighteenth, nineteenth, twentieth, twenty-first, twenty-second, twenty-third, twenty-fourth, twenty-fifth, twenty-sixth, twenty-seventh, twenty-eighth, twenty-ninth, thirtieth, thirty-first, thirty-second, thirty-third, thirty-fourth, thirty-fifth, thirty-sixth, thirty-seventh, thirty-eighth, thirty-ninth, or fortieth tag subset of the tag set;

[0118] A mapping relationship between the tag set and the region set is constructed:

[0119] ;

[0120] The tag is taken as input text , and a visual language model is used to obtain output text :

[0121] ,

[0122] In the formula, is the input text, is the forestry image, is the output text;

[0123] The output text is displayed on the bounding box as a forestry image with visual cues.

[0124] The embodiment improves the efficiency of displaying the output text on the bounding box by constructing a mapping relationship between the tag set and the region set.

[0125] The large language model can use commercial LLaVA, GPT, and DeepSeek, refer to Table 1-Comparison of forestry image generation method with visual cues and traditional advanced VLMs performance table, and Table 2-Comparison of precision of the method to traditional advanced models under forestry data set, in the table, LLaVA-1.5+IAVP, GPT-4V+IAVP, and DeepSeek+IAVP are forestry image generation methods with visual cues using LLaVA, GPT, and DeepSeek as large language models, respectively. It is not difficult for those skilled in the art to see that the benchmark model has a significant improvement in accuracy after being improved by the method provided in the embodiment. Tests show that the method can stably improve the performance of the model on different types of visual language models.

[0126] The three embodiments in Table 2 show high precision on specific forestry data sets: to meet the special needs of forestry scenarios, the method integrates dynamic instance perception and semantic enhancement technology in depth, and performs well in professional tasks such as forest resource monitoring. Especially in the identification and management of economic tree species such as Catalpa bungei, the system has achieved significant improvement in identification accuracy and work efficiency, providing reliable technical support for modern intelligent management of forestry.

[0127] Table 1. Performance Comparison of Forestry Image Generation Methods with Visual Cueing and Traditional Advanced VLMs

[0128]

[0129] Table 2. Comparison of accuracy feedback between the original method and traditional advanced models on forestry datasets.

[0130]

[0131] refer to Figure 2 , Figure 2 This reveals the process of this method from another perspective. Figure 2 The flowchart shown includes: Large Language Model; Extraction; Generation; Filtering.

[0132] Input Text: How many grapes are there in that small bunch below?

[0133] Image input; (The image shows a plate of fruit, including apples, oranges, and grapes).

[0134] Example: small bunches of grapes; object detection; object detection model;

[0135] Label generation (the image shows a plate of fruit, with the grape area marked); Visual-Language Model; Precise Detection; Interleaved Prompt; Scene Understanding.

[0136] Figure 3 This demonstrates how to generate bounding boxes in fruit images based on user-generated natural language descriptions.

[0137] Figure 4 and Figure 5 The method provided in this embodiment is demonstrated to identify specific trees and provide visual cues in forestry images under different weather conditions.

[0138] Example 3:

[0139] The embodiment provides a computer device, including a processor and a memory connected with the processor, and the memory stores a computer program, when the computer program is executed by the processor, the steps of the forestry image generation method with visual prompts provided in the first or second embodiment are executed.

[0140] The computer device can be a server or an electronic terminal, as one of the embodiments, referring to Figure 6 The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is used to store the data obtained and generated in the forestry image generation method with visual prompts. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through network connection. The computer program is executed by the processor to implement the forestry image generation method with visual prompts provided in the first or second embodiment.

[0141] Those skilled in the art can understand that Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0142] The computer device provided by the embodiment has the same technical effects as the first or second embodiment, which will not be described here.

[0143] Embodiment four:

[0144] The embodiment provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the steps of the forestry image generation method with visual prompts provided in the first or second embodiment.

[0145] The computer readable storage medium provided by the embodiment has the same technical effects as the first or second embodiment, which will not be described here.

[0146] Embodiment five:

[0147] The embodiment provides a computer program product, which stores a computer program, and the computer program is executed by a processor to realize the steps of the forestry image generation method with visual cues provided in the embodiment one or the embodiment two. The computer program product provided in the embodiment can be transmitted, distributed and downloaded in the form of signals through the Internet.

[0148] The computer program product provided in the embodiment has the same technical effects as the embodiment one or two, and details are not described herein.

[0149] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. In addition, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0150] The present application is described with reference to flowcharts and / or block diagrams according to the methods, devices (systems) and computer program products of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device implemented in the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one or more flows and / or blocks. Figure 1 The device that implements the functions specified in one or more flows and / or blocks.

[0151] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including instruction devices, which implement the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one or more flows and / or blocks. Figure 1 The device that implements the functions specified in one or more flows and / or blocks.

[0152] These computer program instructions can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one or more flows and / or blocks. Figure 1 The device that implements the functions specified in one or more flows and / or blocks.

[0153] In addition, the terms "first", "second", etc. are used only for descriptive purposes and should not be construed as indicating or implying relative importance or an indicated number of technical features. Thus, the features defined with "first", "second", etc. can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified and limited, the term "mounting", "connection", "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integral connection; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through intermediate medium, or internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0154] In the description of the present application, it should be noted that, unless otherwise specified and limited, the terms "mounting", "connection", "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integral connection; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through intermediate medium, or internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0155] The above is only the preferred embodiment of the present application, it should be pointed out that, for those skilled in the art, without departing from the technical principles of the present application, a number of improvements and modifications can be made, these improvements and modifications should also be considered as the protection scope of the present application.

Claims

1. A method for generating forestry images with visual cues, characterized in that, include, Obtain the input language; The input language is identified by a large language model that is pre-linked with a knowledge enhancement module of forestry concept system, and multiple standard instance descriptions that conform to forestry survey specifications are obtained. All standard instance descriptions are prioritized according to their relevance to the intent of the input language. Image regions are delineated from the forestry images based on each standard instance description; The image regions are visually cued and labeled according to the priority described in each standard instance to obtain a forestry image with visual cues.

2. The forestry image generation method with visual cues according to claim 1, characterized in that, The forestry images are pre-divided based on forest area categories to obtain partitioned images, and each partitioned image is labeled with the corresponding forest area category; Visual cues were labeled for each partition image, and the partition images with visual cues were merged to obtain a forestry image with visual cues. The prioritization of all standard instance descriptions according to their relevance to the intent of the input language includes, All standard instance descriptions are prioritized based on their relevance to the intent of the input language and their relevance to the language concepts under the forest area category.

3. The forestry image generation method with visual cues according to claim 1, characterized in that, include, The process involves identifying the input language using a large language model pre-associated with a forestry concept system knowledge enhancement module, resulting in multiple standard instance descriptions conforming to forestry survey specifications, including: The large language model extracts noun phrases from the input language and organizes these noun phrases into a candidate instance set C = {c1, c2, ..., cn}. The formula for obtaining the candidate instance set is as follows: , In the formula, For large language models, Indicates the extraction action. Let c be the input language, n be the number of candidate instances, C be the candidate instance set, and c1 and c2 represent the first and second candidate instances in the candidate instance set. The large language model utilizes a forestry concept system knowledge enhancement module to semantically expand and contextually associate the candidate instances, generating multiple standard instance descriptions, which are then organized into a first instance set. The first instance set is used to obtain forestry images with visual cues, and the instance set is obtained as follows: ; In the formula, Let m be the first set of instances, and m be the number of standard instance descriptions. and This is represented as the first and second standard instance descriptions. This indicates the generation of an action.

4. The forestry image generation method with visual cues according to claim 3, characterized in that, The process of recognizing the input language through a large language model pre-associated with a forestry concept system knowledge enhancement module, and obtaining multiple standard instance descriptions conforming to forestry survey specifications, also includes... The large language model obtains the query intent based on the input language, and filters the first instance set according to the query intent to obtain the second instance set. The second set of instances is used to obtain forestry images with visual cues; In the formula, For the second set of instances, The number of standard instance descriptions in the second instance set.

5. The forestry image generation method with visual cues according to claim 4, characterized in that, The step of dividing the image region from the forestry image according to each standard instance description includes, Using image processing tools, the forestry image is divided into multiple image regions based on the second instance set. Each standard instance description corresponds to one or more image regions; Image regions corresponding to the same standard instance are integrated into a subset of regions: , In the formula, The serial number described for the standard instance. For the corresponding number A subset of regions described by a standard instance, For image region, This is the index of the image region within the subset. This represents the number of image regions within the subset. For the corresponding number The first, second... of the regions described in the standard instance subset Image regions; Integrate the subsets of regions corresponding to different standard instances into a region set: , In the formula, For a set of regions, The number of region subsets, For the 1st, 2nd... in the set of regions A subset of regions; Construct the mapping relationship between the second set of instances and the set of regions: ; The step of visually cuing the image regions according to the priority described in each standard instance to obtain a forestry image with visual cues includes: An image processing tool is acquired, which generates a bounding box covering the image region on the forestry image using the following formula: , In the formula, For forestry images with bounding boxes, For the input image, This refers to the operation of image processing tools.

6. The forestry image generation method with visual cues according to claim 5, characterized in that, The image processing tool acquires forestry images with bounding boxes using the following formula, including: Sort all bounding boxes according to pixel area; Generate bounding boxes in descending order of pixel area; In response to the overlap between a bounding box generated later in the generation order and a bounding box generated earlier in the generation order, the position of the bounding box generated later in the generation order is adjusted until they no longer overlap. The bounding boxes corresponding to different standard instances are displayed with separate color shading.

7. The forestry image generation method with visual cues according to claim 5, characterized in that, The step of visually cuing the image region according to the priority described in each standard instance to obtain a forestry image with visual cues also includes: The large language model is used to generate tag information corresponding to all standard instance descriptions; Organize the aforementioned tag information into a tag set: ,in Corresponding to the The first standard instance describes the corresponding region subset. Labels for each image region In the formula, For a set of tags, For the 1st, 2nd... in the tag set A subset of tags, For the first tag in the tag set A subset of tags; Construct a mapping relationship between the tag set and the region set: ; Using the label as input text Use a visual language model to obtain the output text : , In the formula, For input text, For forestry images, To output text; The output text is displayed on the bounding box as a forestry image with visual cues.

8. A computer device, characterized in that, The method includes a processor and a memory connected to the processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the forestry image generation method with visual cues as described in any one of claims 1 to 7 are performed.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the forestry image generation method with visual cues as described in any one of claims 1 to 7.

10. A computer program product comprising a computer program / instructions, characterized in that, When executed by a processor, the computer program / instructions implement the steps of the forestry image generation method with visual cues as described in any one of claims 1 to 7.