Multi-modal data generation and fine adjustment method for domain image

By generating multimodal data combining images and text and fine-tuning the model, the general visual multimodal model's poor performance in professional field images is solved, and its recognition and question-and-answer capabilities in field images are improved.

CN120046117AActive Publication Date: 2025-05-27BEIJING ZHONGKE STRON CLOUD INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510511519.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-27
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing general visual multimodal model performs poorly in professional field images, mainly due to the lack of image data in professional field images involved in the training process, resulting in insufficient capabilities in these fields.

Method used

By using existing multimodal models and large language models to process field images, combining multi-layer extraction, multimodal data combining images and text are generated, and a general visual multimodal model is fine-tuned based on this.

Benefits of technology

Improves the recognition and question-and-answer capabilities of general visual multimodal models in domain images, allowing them to process image data in professional fields more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046117A_ABST
    Figure CN120046117A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal data generation and fine tuning method for a domain image, and belongs to the technical field of artificial intelligence, and the method comprises the steps: carrying out the primary processing of the domain image through a visual multi-modal model, obtaining a general description, carrying out the natural language conversion of a set label, obtaining a conversion text, and carrying out the fine tuning of the conversion text; carrying out semantic integration and expansion on the image description and the converted text based on a large language model to generate an initial description; performing multi-layer extraction on the domain image according to description dimensions to obtain layer description of each layer, and generating comprehensive description in combination with the initial description; matching the domain image with the corresponding comprehensive description to obtain multi-modal data; and performing fine tuning processing on the visual multi-modal model by using the multi-modal data to obtain an optimized multi-modal model. And the identification and question-answering capabilities in the domain image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method for generating and fine-tuning multi-modal data of domain images. Background Art

[0002] Visual multi-modal models (such as gpt-4o / qwen-vl, etc.) can combine text and image information to achieve vision-based question-answering functions. Compared with traditional image classification and image recognition methods, such models have a better user experience and stronger image understanding ability. However, the current general visual multi-modal models still perform unsatisfactorily on domain images. This is mainly because the professional domain image data involved in the training process of general models is less, resulting in insufficient capabilities in these fields.

[0003] For example, in professional fields such as biological species classification, general visual multi-modal models may be difficult to provide sufficiently accurate answers. And in each professional field, although there are a large number of domain image data sets (such as image classification, image detection data sets), these data are usually unimodal, that is, the images only correspond to numerical labels and lack detailed text descriptions, making it difficult to directly use them for the training of visual multi-modal models.

[0004] In summary, how to convert unimodal domain image data into image-text multi-modal data that can be used for multi-modal model training, and fine-tune the model with these data to improve its capabilities in professional fields is an urgent problem to be solved.

[0005] Therefore, the present invention proposes a method for generating and fine-tuning multi-modal data of domain images. Summary of the Invention

[0006] The present invention provides a method for generating and fine-tuning multi-modal data of domain images, which is used to process domain images by using existing multi-modal models and large language models, and combine multi-layer extraction to generate multi-modal data combining images and text, and fine-tune the general visual multi-modal model based on this, so as to improve its recognition and question-answering capabilities in domain images.

[0007] The present invention provides a method for generating and fine-tuning multi-modal data of domain images, including: Step 1: Use a visual multi-modal model to preliminarily process a domain image to obtain a general description, where the general description includes: an image description and a set label for the image description; Step 2: Perform natural language conversion on the set label to obtain a conversion text, and semantically integrate and expand the image description and the conversion text based on a large language model to generate an initial description; Step 3: Perform multi-layer extraction on the domain image according to the description dimension to obtain the layer description of each layer, and generate a comprehensive description in combination with the initial description; Step 4: Pair the domain image with the corresponding comprehensive description to obtain multi-modal data; Step 5: Use the multi-modal data to perform fine-tuning processing on the visual multi-modal model to obtain an optimized multi-modal model.

[0008] Preferably, performing multi-layer extraction on the domain image according to the description dimension to obtain the layer description of each layer includes: Determine the domain type of the domain image, and match the description dimension and the description definition based on each description dimension from the type-dimension comparison table; Respectively match the dimension extraction model from the type-definition-recognition comparison table according to the description definition, and perform layer extraction on the domain image to obtain the extraction layer under the corresponding description dimension, and then perform meaning analysis on the extraction layer to obtain the layer description.

[0009] Preferably, performing meaning analysis on the extraction layer to obtain the layer description includes: Retrieve the process log of the layer extraction of the domain image by the dimension extraction model, and respectively determine the extraction key coefficient of each position point in the corresponding extraction layer; Divide the extraction key coefficient according to the extraction threshold matched with the corresponding description dimension, and regard the surface composed of the position points greater than or equal to the extraction threshold as the first layer, and regard the surface composed of the position points less than the extraction threshold as the second layer; Perform global feature analysis on the corresponding extraction layer to obtain the first feature. At the same time, perform global feature analysis on the corresponding first layer and second layer respectively to obtain the corresponding second feature and third feature; Perform natural language processing on the first feature, second feature and third feature respectively to obtain the corresponding first description, second description and third description; Establish a first difference function between the first description and the second description, a second difference function between the first description and the third description, and a fusion function between the second description and the third description; Obtain supplementary features according to the first difference function, second difference function and fusion function; Perform feature meaning analysis on the first feature and the supplementary features to obtain the corresponding layer description.

[0010] Preferably, generating a comprehensive description in combination with the initial description includes: If there is only 1 description dimension under the domain type, at this time, regard the corresponding layer description as the description to be analyzed; If there are multiple description dimensions under the described domain type, at this time, perform description fusion on the layer descriptions under all description dimensions to obtain the description to be analyzed; Generate a comprehensive description based on the description to be analyzed and in combination with the initial description.

[0011] Preferably, performing description fusion on the layer descriptions under all description dimensions to obtain the description to be analyzed includes: Extract the quantity vocabulary and category vocabulary in each layer description to construct a description vector; Based on deep learning technology, mine the interest definitions of the category vocabulary under each description dimension, and obtain the interested vocabulary under the corresponding category vocabulary, and then construct an extended vector; Determine the interested type of each category vocabulary in the extended vector for the interested vocabulary, and count the number of occurrences of the same interested type in the extended vector; Screen the interested type with the largest statistical count, and match the corresponding attention mechanism from the type-mechanism comparison table; Perform interested description embedding on the layer descriptions of all extraction layers based on the attention mechanism under each description layer, and then perform description fusion to obtain the description to be analyzed.

[0012] Preferably, obtaining the interested vocabulary under the corresponding category vocabulary includes: Respectively use the category vocabulary under the description dimension as an index, and in combination with deep learning technology, mine the strongly related definitions based on each index under the corresponding description dimension from the specified database; Perform clustering analysis on all the strongly related definitions under each index to obtain a definition cluster, and regard the referring definition of the corresponding definition cluster as the interest definition, where the number of the definition clusters is at least 1.

[0013] Preferably, using the multimodal data to perform fine-tuning processing on the visual multimodal model includes: Convert the multimodal data into the current input sequence matching the visual multimodal model; Convert the general description of the domain image into the original input sequence matching the visual multimodal model; Compare and analyze the original input sequence and the current input sequence of the same domain image, determine the difference sequence and the first number Nd of the difference sequence, and according to Construct a division window according to the obtained second number, where represents the ceiling symbol; represents randomly extracting 1 positive integer from the range ; Lock the start position where the difference sequence first appears and the end position where the difference sequence last appears in the current input sequence to obtain a position sequence segment; If the segment window of the position sequence segment is the same size as the division window, keep the current input sequence unchanged; If not, divide the position sequence segment in sequence according to the division window at the starting position to obtain several combined sequences, and combine each combined sequence, the sequence before the starting position, and the sequence after the ending position respectively to obtain a new sequence; Perform natural language processing on the new sequence. If the semantics of the new sequence are similar to those of the current input sequence, keep the corresponding new sequence; Otherwise, screen the positions where the differential sequences appear for the second time, and combine with the ending position where the differential sequences appear for the last time to re-judge the newly obtained sequences subsequently; Train the visual multi-modal model based on all the retained new sequences and all the current input sequences.

[0014] Preferably, based on a large language model, semantically integrate and expand the image description and the converted text to generate an initial description, including: Based on a large language model, find relevant concepts in the image description and the converted text and establish corresponding relationships; Based on the corresponding relationships, fuse the complementary information in the image description and the converted text; Utilize an external knowledge graph to search for newly introduced concepts and knowledge related to the fused semantics, and combine with the background knowledge of the domain images to determine the depth and breadth of the semantics, generating an integrated and expanded text; Re-express the integrated and expanded text in natural language to obtain the initial description.

[0015] Compared with the prior art, the beneficial effects of the present application are as follows: By using existing multi-modal models and large language models to process domain images, and combining multi-layer extraction to generate multi-modal data combining images and texts, and based on this, fine-tuning a general visual multi-modal model, thereby improving its recognition and question-answering capabilities in domain images.

[0016] Other features and advantages of the present invention will be described in the subsequent specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the written specification and the drawings.

[0017] The technical solutions of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings

[0018] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings: Figure 1 It is a flowchart of a method for generating and fine-tuning multimodal data of a domain image in an embodiment of the present invention. Detailed implementation manners

[0019] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0020] The present invention provides a method for generating and fine-tuning multimodal data of a domain image, as Figure 1 shown, including: Step 1: Use a visual multimodal model to perform preliminary processing on the domain image to obtain a general description, where the general description includes: an image description and a set label for the image description; Step 2: Perform natural language conversion on the set label to obtain a conversion text, and based on a large language model, semantically integrate and expand the image description and the conversion text to generate an initial description; Step 3: Perform multi-layer extraction on the domain image according to the description dimension to obtain a layer description for each layer, and combine the initial description to generate a comprehensive description; Step 4: Pair the domain image with the corresponding comprehensive description to obtain multimodal data; Step 5: Use the multimodal data to perform fine-tuning processing on the visual multimodal model to obtain an optimized multimodal model.

[0021] In this embodiment, the visual multimodal model can be the Qwen2.5-VL-7B-Instruc model. To obtain labeled image data within the domain, for example, image classification data of fungi.

[0022] Send the picture and the prompt "Please describe this picture" to the visual multimodal model to obtain an image description, such as "There is a mushroom in this picture".

[0023] In this embodiment, the set label of the image description can be a numerical or character label.

[0024] In this embodiment, the set tags are subjected to natural language conversion to obtain converted text. For example, in an image classification dataset, the tags of images are usually numbers (e.g., the tag "1" represents a cat, "2" represents a dog, etc.). By designing natural language templates, these numerical tags are converted into meaningful text descriptions. For example, the tag "1" can be converted to "The species in this image is a cat", and the tag "2" can be converted to "This is a dog". This conversion helps to turn simple numerical tags into understandable text information, facilitating the combination with image descriptions.

[0025] In this embodiment, the tuning framework for the model generally includes: a Vision Encoder, a Large Language Model (LLM), and an Adapter.

[0026] The Vision Encoder is used to extract features from images. The Adapter serves as a bridge to convert image features into the word embedding space, thus facilitating the LLM's interpretation of the output of the Vision Encoder. The Adapter is usually designed as a lightweight neural network structure, such as a fully connected network, to ensure efficient multimodal integration. Subsequently, the LLM processes the data combined with text and image features to generate the expected text response.

[0027] Loss function for model training: The goal of training is to minimize the gap between the generated text and the true text. Usually, the Cross-Entropy Loss is used to calculate the difference between the probability of the generated text and the true label. Specifically, the model is trained by maximizing the probability of the correct text. The loss function Lgeneration can be expressed as: ; where N is the number of samples in a training batch, i represents the i-th sample in the training batch, T is the length of the valid output sequence y i of sample i, xi is the feature vector of the image in training batch sample i, INSTi is the text instruction corresponding to the image in sample i, for example, "Please describe this image". is the true label at time step t of the output sequence of sample i. is the generated sequence of sample i before time step t. indicates that the output sequence time step t of sample i is the true label probability.

[0028] During the training process, the model will continuously adjust the parameters through Backpropagation to minimize the loss function.

[0029] At each iteration, the model generates an output response based on the input image and instruction text, then compares it with the true annotation text, calculates the loss, and updates the model parameters.

[0030] After training, the model is evaluated on the validation set. The evaluation metrics include: Custom metric: Evaluate the accuracy of the model's prediction of image labels. For example, when inputting a picture of a certain type of mushroom, output a prediction of whether it is edible.

[0031] Since the label information of the image is incorporated into the image description during the data generation stage, the accuracy of label prediction should be evaluated during prediction. The method can be designed according to the specific labels. For example, keyword matching can be used to determine whether the prediction is accurate.

[0032] Common evaluation metrics for text generation tasks include BLEU or ROUGE, etc.

[0033] The evaluation process includes: obtaining the sample data of the validation set, inputting each sample data (image and text instruction) into the trained multi-modal model to get the predicted text output, and then evaluating using the custom metric and common metrics based on the predicted output and the correct output given in the sample.

[0034] In this embodiment, the generated domain multi-modal data is used to fine-tune the general vision multi-modal model through supervised learning.

[0035] Optimize the model parameters to make it more suitable for domain image data, thereby enhancing its image understanding and question-answering capabilities in a specific domain.

[0036] In this embodiment, it is necessary to analyze the domain image in three dimensions. For example, the environmental scene dimension of the image (such as grassland, sky), the animal scene dimension in the image (such as rabbit), and the human scene dimension in the image (such as person). Then, based on these three dimensions, layer extraction is performed on the image to obtain separate images corresponding to each dimension. For example, the extraction layer obtained for the environmental scene dimension only contains grassland, sky, etc. (excluding rabbits and people), and so on, to obtain different extraction layers.

[0037] In this embodiment, after obtaining the extraction layer, analyzing the image of the corresponding layer can obtain the layer description.

[0038] In this embodiment, the comprehensive description is obtained through the comprehensive processing of all layer descriptions and the initial description.

[0039] In this embodiment, pairing refers to correlating the comprehensive description with the domain image. The multi-modal data refers to the comprehensive description of the domain image.

[0040] In this embodiment, since the description richness of the multimodal data is higher than the original general description, that is, the samples are optimized. Therefore, by continuing to train the model with the multimodal data, the fine-tuning of the model accuracy can be achieved, and then the optimized multimodal model can be obtained.

[0041] The beneficial effects of the above technical solution are as follows: By using the existing multimodal model and large language model to process domain images, and combining multi-layer extraction to generate multimodal data combining images and texts, and based on this, fine-tuning the general vision multimodal model, so as to improve its recognition and question-answering capabilities in domain images.

[0042] The present invention provides a method for generating and fine-tuning multimodal data of domain images, which performs multi-layer extraction on the domain images according to the description dimension to obtain the layer description of each layer, including: Determine the domain type of the domain image, and match the description dimension and the description definition based on each description dimension from the type-dimension comparison table; Respectively match the dimension extraction model from the type-definition-recognition comparison table according to the description definition, perform layer extraction on the domain image to obtain the extraction layer under the corresponding description dimension, and then perform meaning analysis on the extraction layer to obtain the layer description.

[0043] In this embodiment, the domain type refers to outdoor scene type, indoor scene type, bungee jumping scene type, etc., and different types are involved for different scenes.

[0044] In this embodiment, the type-dimension comparison table contains different domain types and the set description dimensions and description definitions based on this type, which are stored in advance and can be directly matched.

[0045] In this embodiment, the type-definition-recognition comparison table contains dimension extraction models based on different description definitions under different domain types, which are stored in advance and can be directly matched. Moreover, the dimension extraction model is trained on a neural network model with the original image under the corresponding description dimension and the scene extraction result under this dimension as samples.

[0046] The specific content included in the type-definition-recognition comparison table is as follows:

[0047] Suppose the input image is a photo of a forest: Determine that the domain type is an outdoor scene; Search for the description dimension and model. According to the type-definition-recognition comparison table, the following information is found: Description dimension 1: Color of the sky Description definition: Extract the main color of the sky area Dimension extraction model: Color segmentation model (based on CNN) Description dimension 2: Texture of the ground Description definition: Identify the texture features of the ground area Dimension extraction model: Texture classification model (based on ResNet) Description dimension 3: Types of plants Description definition: Identify the main plant species in the plant area Dimension extraction model: Plant classification model (based on YOLOv5) Execute extraction: Process the image using the corresponding model: Color of the sky: The model outputs "blue".

[0048] Texture of the ground: The model outputs "grassland".

[0049] Types of plants: The model outputs "pine trees".

[0050] Generate layer description: Fuse the extraction results to obtain the final layer description: "This is a sunny forest scene with a blue sky, the ground covered with grassland, and the main plants being pine trees."

[0051] In this embodiment, after using the model to perform scene extraction under the corresponding description dimension, the extraction layer (the scene extraction result based on the model) can be obtained.

[0052] In this embodiment, the meaning analysis is performed by dividing the threshold of the position points and then fusing the features and descriptions.

[0053] The beneficial effects of the above technical solution are: Obtain the description definition and the relevant dimension extraction model from the two comparison tables in sequence, thereby obtaining the extraction layer, and obtaining the layer description through meaning analysis, providing a basis for subsequent obtaining of the comprehensive description.

[0054] The present invention provides a method for generating and fine-tuning multi-modal data of domain images, and performing meaning analysis on the extraction layer to obtain a layer description, including: Retrieve the process log of the layer extraction of the domain image by the dimension extraction model, and respectively determine the extraction key coefficients of each position point in the corresponding extraction layer; Divide the extraction key coefficients according to the extraction threshold matching the corresponding description dimension, and regard the surface composed of the position points greater than or equal to the extraction threshold as the first layer, and regard the surface composed of the position points less than the extraction threshold as the second layer; Perform global feature analysis on the corresponding extraction layer to obtain the first feature. At the same time, perform global feature analysis on the corresponding first layer and second layer respectively to obtain the corresponding second feature and third feature; Perform natural language processing on the first feature, second feature, and third feature respectively to obtain the corresponding first description, second description, and third description; Establish a first difference function between the first description and the second description, a second difference function between the first description and the third description, and a fusion function between the second description and the third description; Obtain supplementary features according to the first difference function, second difference function, and fusion function; Perform feature meaning analysis on the first feature and the supplementary features to obtain the corresponding layer descriptions.

[0055] In this embodiment, the process log refers to the relevant data generated during the process of the model extracting layers from the image, which can be directly captured. Specifically: Before the model starts processing the image, use the time function in the programming language to record the start time. For example, in Python, the time.time( ) function can be used to obtain the current timestamp as the start time start_time.

[0056] At the start of each pixel extraction operation, record the start time pixel_start_time of the extraction of this pixel point. Similarly, the time.time( ) function can be used.

[0057] When the pixel extraction operation ends, record the end time pixel_end_time. Calculate the extraction time of this pixel point by pixel_end_time - pixel_start_time and record it in the log.

[0058] Set a counter for each pixel point with an initial value of 0. A two-dimensional array extraction_count with the same size as the image pixel point matrix can be used to record the extraction times of each pixel point. Each element of the array corresponds to a pixel point in the image.

[0059] Whenever an extraction operation is performed on a pixel point, increment the corresponding counter value by 1. For example, when performing extraction on the pixel point with coordinates (x, y), execute extraction_count[x][y]+=1.

[0060] To create a log file, you can use the open( ) function in Python to open a file in write mode. For example, log_file = open('process_log.txt', 'w').

[0061] During the layer extraction process of the model for an image, after completing the extraction of a pixel point and obtaining its extraction time and updated extraction count, write this information to the log file in a certain format. For example, write it in the format of pixel point coordinates, extraction time, extraction count. For instance, (10,20),0.001,3 means that the extraction time of the pixel point at coordinates (10,20) is 0.001 seconds and the extraction count is 3 times.

[0062] After the model finishes processing the entire image, close the log file to ensure that all log information is correctly saved. Use log_file.close( ) to close the file.

[0063] The process log contains the extraction time and extraction count for each pixel point in the image to determine the extraction focus coefficient for each position point in the corresponding extraction layer. Different image feature extraction algorithms handle pixel points differently. For example, the SIFT (Scale-Invariant Feature Transform) algorithm needs to detect stable feature points in the image and calculate their local features. The computational amount for each pixel point is relatively large, the extraction time is relatively long, and it may sample and calculate certain key pixel points multiple times to ensure the accuracy of the features.

[0064] Among them, the extraction focus coefficient = , where St and Rc respectively represent the extraction time length and extraction count of the corresponding position point; T0 and R0 respectively represent the preset time length and preset count; In this embodiment, the value of the extraction threshold is 1.

[0065] In this embodiment, the first layer is the surface composed of the position points that satisfy being greater than or equal to the extraction threshold; The second layer is the surface composed of the position points that satisfy being less than the extraction threshold.

[0066] In this embodiment, the first feature, the second feature, and the third feature are obtained based on a feature analysis model (a commonly used convolutional neural network (CNN) model in the field of computer vision. It extracts different levels and types of features from the original image data through operations such as multi-layer convolution and pooling on the image. For example, the output of the fully connected layer of a convolutional neural network (CNN) is a feature vector. For instance, after a trained ResNet model extracts features from an image of a cat, it outputs a 1000-dimensional vector, and each value in the vector represents the response intensity of the image in the corresponding feature dimension, which may include numerical representations of relevant features such as the shape of the cat's ears, the color of its eyes, and the texture of its fur). This is prior art. Image features refer to a set of attributes that can characterize the characteristics or content of an image, mainly including image natural features (such as brightness, color, texture, etc.). Furthermore, relevant descriptions can be directly obtained through natural language processing. Specifically, natural language processing technology analyzes the identified image features. It is based on a large amount of training data and complex algorithm models and can convert these feature information into natural language. For example, it will generate a concise natural language description such as "There is a mushroom in the picture" based on the natural features such as the shape and color of the mushroom in the image.

[0067] In this embodiment, the first difference function = G (the first description, the second description), the second difference function = G (the first description, the third description), and the fusion function = R (the second description, the third description). It should be noted that the fusion function is obtained by processing based on a relevant fusion model, and the fusion model is trained on a neural network model with different descriptions and the results of description combinations as samples.

[0068] The difference function is obtained by analyzing based on a difference model, and the difference model is trained on a neural network model with different descriptions and description differences as samples.

[0069] For example, the first description is: description u1, description u2, description u3, description u4, and the second description is description u1, description u2. At this time, the result obtained by the first difference function is: description u3, description u4.

[0070] In this embodiment, the supplementary feature is regarded as the image feature under the inconsistent descriptions corresponding to the difference function and the fusion function. Specifically, the inconsistent descriptions are mapped to the image for comparison, and then the features of the image information under the inconsistent descriptions are obtained.

[0071] In this embodiment, the first feature and the supplementary feature are input into the meaning analysis model to obtain a layer description, and the meaning analysis model is trained on a neural network model with different feature combinations and the analysis results based on these combinations as samples. For example, the layer description for the environmental scene is: The environmental beautification makes people feel relaxed and happy.

[0072] The beneficial effects of the above technical solution are as follows: The extraction focus coefficient of each position point is determined based on the process log, and then coefficient division is achieved through the extraction threshold, effectively determining the features of different layers. Through the analysis of the features of the three layers and natural language processing, layer descriptions are obtained, ensuring the accuracy of description acquisition.

[0073] The present invention provides a method for generating and fine-tuning multimodal data of domain images, which generates a comprehensive description in combination with the initial description, including: If there is only 1 description dimension under the domain type, at this time, the corresponding layer description is regarded as the description to be analyzed; If there are multiple description dimensions under the domain type, at this time, the layer descriptions under all description dimensions are subjected to description fusion to obtain the description to be analyzed; Based on the description to be analyzed and in combination with the initial description, a comprehensive description is generated.

[0074] In this embodiment, the comprehensive description refers to performing a description union process on the description to be analyzed and the initial description to obtain the comprehensive description.

[0075] The beneficial effects of the above technical solution are as follows: By performing a quantitative analysis on the dimensions, the rationality of obtaining the description to be analyzed is realized, and then a comprehensive description is obtained.

[0076] The present invention provides a method for generating and fine-tuning multimodal data of domain images, which performs description fusion on the layer descriptions under all description dimensions to obtain the description to be analyzed, including: Extract the quantity words and category words in each layer description to construct a description vector; Based on deep learning technology, mine the interest definition of the category words under each description dimension, and obtain the interesting words under the corresponding category words, and then construct an extended vector; Determine the interesting type of each category word in the extended vector for the interesting words, and count the number of occurrences of the same interesting type in the extended vector; Select the interesting type with the largest statistical count, and match the corresponding attention mechanism from the type-mechanism look-up table; Based on the attention mechanism under each description layer, perform interesting description embedding on the layer descriptions of all extraction layers, and then perform description fusion to obtain the description to be analyzed.

[0077] In this embodiment, the description vector = {category words in the layer description, quantity words under the corresponding category words}.

[0078] In this embodiment, the extended vector = {interesting words under the category words}, and the extended vector includes the words involved in the corresponding layer description.

[0079] In this embodiment, under the category term 1: interesting term 11, interesting term 12; Under the category term 2: interesting term 21, sensing area term 22, interesting term 23; The interesting type A1 includes: interesting term 11, interesting term 21, sensing area term 22; The interesting type A2 includes: interesting term 12, interesting term 23; At this time, the number of terms under the interesting type A1 obtained by statistics is 3, and the number of terms under the interesting type A2 obtained by statistics is 2.

[0080] In this embodiment, the type-mechanism correspondence table includes different interesting types and the attention mechanism for this type, and the attention mechanism is determined in advance for this type.

[0081] In this embodiment, the interesting description embedding means that since the attention mechanism will make the model focus on the key parts of the input text, by calculating the attention weights at each position, it is determined which information is more important for generating the interesting description. Then, according to these weights, the features of the input text are integrated, and the information related to the interesting content is effectively embedded into a low-dimensional vector space to form the embedded representation of the interesting description.

[0082] In this embodiment, the description fusion is implemented based on the fusion description model, and the fusion description model is trained with the neural network model using the interesting description embeddings of the description layers under different description dimensions and the related fusion results as samples.

[0083] The beneficial effects of the above technical solutions are: by constructing the description vector and mining the interest definition, an augmented vector is constructed to provide a basis for determining the largest interesting type subsequently, and the attention mechanism is obtained to realize the reasonable acquisition of the interesting description embedding, ensuring the diversity and integrity of the information in the description to be analyzed.

[0084] The present invention provides a method for generating and fine-tuning multimodal data of domain images, obtaining interesting terms under corresponding category terms, including: Respectively using the category terms under the description dimension as indexes, and combining deep learning techniques to mine the strongly related definitions based on each index under the corresponding description dimension from the specified database; Performing clustering analysis on all the strongly related definitions under each index to obtain definition clusters, and regarding the representative definitions of the corresponding definition clusters as interest definitions, where the number of the definition clusters is at least 1.

[0085] In this embodiment, the specified database pre-stores relevant field text information including interesting vocabulary under different categories, and then strong relevant definitions can be obtained by mining the text information.

[0086] In this embodiment, the strong relevant definition means that the correlation coefficient between the vocabulary in the specified database and the corresponding index is greater than 0.9, where the calculation method of the correlation coefficient is existing.

[0087] In this embodiment, if there are 10 strong relevant definitions under an index, then clustering analysis is performed on these 10 strong relevant definitions to obtain the existing clusters, which are regarded as definition clusters.

[0088] The beneficial effects of the above technical solution are: using category vocabulary as the index and combining deep learning technology to mine strong relevant definitions, and then effectively obtaining interest definitions through clustering analysis.

[0089] The present invention provides a method for generating and fine-tuning multimodal data of domain images, and uses the multimodal data to perform fine-tuning processing on the visual multimodal model, including: Converting the multimodal data into a current input sequence matching the visual multimodal model; Converting the general description of the domain image into an original input sequence matching the visual multimodal model; Comparing and analyzing the original input sequence and the current input sequence of the same domain image to determine the difference sequence and the first number Nd of the difference sequence, and according to Constructing a division window based on the obtained second number, where represents the ceiling symbol; represents randomly extracting 1 positive integer from the range ; Locking the start position where the difference sequence first appears and the end position where the difference sequence last appears in the current input sequence to obtain a position sequence segment; If the segment window of the position sequence segment is the same size as the division window, keep the current input sequence unchanged; If not, sequentially divide the position sequence segment according to the division window at the start position to obtain several combined sequences, and respectively combine each combined sequence, the sequence before the start position, and the sequence after the end position to obtain a new sequence; Performing natural language processing on the new sequence, and if the semantics of the new sequence is similar to the semantics of the current input sequence, retain the corresponding new sequence; Otherwise, screen the position where the difference sequence appears for the second time, and in combination with the end position where the difference sequence appears for the last time, re-judge the subsequent obtained new sequences; The visual multi-modal model is trained based on all the retained new sequences and all the current input sequences.

[0090] In this embodiment, the multi-modal data conversion is to ensure that the input format meets the input format of the visual multi-modal model, so as to obtain the current input sequence, and the general description is similar thereto.

[0091] In this embodiment, for example: the current input sequence = {r1 r2 r3 r4 r5 r6}, the original input sequence = {r1 r6 r7}, at this time, the obtained difference sequence = {r1 - r1 r2 r3 r4 r5 r6 - r6 r7}, at this time, the first number Nd obtained is 5, at this time, the first position where the difference sequence appears for the first time is the position corresponding to r2, the last position where the difference sequence appears for the last time is the position corresponding to r7, the position sequence segment is {r2 r3 r4 r5 r6 - r6 r7}, at this time, if the size of the divided window is 2, the obtained combined sequences are {r2 r3}, {r4 r5}, {r6 - r6 r7} in turn, and the obtained new sequences are: {r1 r2 r3}, {r1, r4 r5}, {r1, r6, r7} in turn.

[0092] In this embodiment, semantic similarity means that after natural language processing, it is found that the semantic similarity between the new sequence and the semantic of the current input sequence is greater than 80% or more, and at this time, it is regarded as similar.

[0093] The beneficial effects of the above technical solutions are: by obtaining multi-modal data and the input sequence of general description, and then determining the position sequence segment through sequence difference comparison and analysis, and combining the size of the divided window to obtain new sequences, providing more samples for the training of the visual multi-modal model, and ensuring the reliability of the samples through semantic similarity analysis, so as to realize the fine-tuning process of the model.

[0094] The present invention provides a method for generating and fine-tuning multi-modal data of domain images, which integrates and expands the semantic meanings of image descriptions and conversion texts based on a large language model to generate an initial description, including: Finding relevant concepts in the image description and conversion text based on the large language model and establishing corresponding relationships; Based on the corresponding relationships, fusing the complementary information in the image description and conversion text; Using an external knowledge graph to find newly introduced concepts and knowledge related to the fused semantics, and combining the background knowledge of the domain image to determine the depth and breadth of the semantics, and generating an integrated and expanded text; Re-expressing the integrated and expanded text in natural language to obtain the initial description.

[0095] In this embodiment, the relevant concepts are all pre-set based on a large prediction model. For example, the image description is: a scene of a little dog lying curled up on the grass, and the converted text is: a little dog lying curled up on the grass. At this time, the image description and the corresponding converted text have a corresponding relationship.

[0096] In this embodiment, semantic fusion is to merge multiple knowledge graphs and unify the entities, concepts, and relationships in different knowledge graphs. The entity alignment problem needs to be solved, that is, to identify the nodes representing the same real-world entity in different knowledge graphs and merge them or establish associations to form a more comprehensive and integrated knowledge graph, thereby obtaining the fusion information.

[0097] In this embodiment, the newly introduced concepts and knowledge refer to the supplementary new concepts and new information. For example, when generating the integrated and extended text, the newly introduced concepts and knowledge found are incorporated into it, such as "Currently, for the treatment of diabetes, in addition to traditional drugs, the newly launched [drug name] has shown good efficacy, and its mechanism of action is [specific mechanism]. In terms of combination therapy, the combination plan of [drug combination] can more effectively control blood sugar levels, but attention should be paid to [precautions]." At the same time, the deepening of semantics by combining the background knowledge of the image is also reflected in the text, such as "When there is a lung shadow in the image, if the shadow is located in [specific lung lobe position] and presents [describe the characteristics of the shadow shape, boundary, etc.], it may be [specific disease], and subsequent [other examination items] need to be combined for further diagnosis." In this way, through the collaborative use of external knowledge graphs and domain image background knowledge, the in-depth mining of semantics and the rich expansion of the text are realized.

[0098] In this embodiment, rephrasing is to make the description of the integrated and extended text more fluent.

[0099] The beneficial effects of the above technical solutions are: By establishing the corresponding relationship between the description and the text through a large language model, and then through subsequent information fusion, graph search, background knowledge assistance, etc., the integrated and extended text is obtained to determine the information integrity of the text, and through the re-expression of natural language, the description is made more process, providing reliable samples for subsequent model training.

[0100] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.

Claims

1. A method for generating and fine-tuning multimodal data of domain images, characterized in that: include: Step 1: Use a visual multimodal model to preliminarily process the domain image to obtain a general description, wherein the general description includes: an image description and a set label for the image description; Step 2: Performing natural language conversion on the set label to obtain converted text, and semantically integrating and expanding the image description and the converted text based on a large language model to generate an initial description; Step 3: Perform multi-layer extraction on the domain image according to the description dimension to obtain a layer description of each layer, and generate a comprehensive description in combination with the initial description; Step 4: Pair the domain image with the corresponding comprehensive description to obtain multimodal data; Step 5: Use the multimodal data to fine-tune the visual multimodal model to obtain an optimized multimodal model.

2. The method for generating and fine-tuning multimodal data of domain images according to claim 1, characterized in that: The domain image is subjected to multi-layer extraction according to the description dimension to obtain a layer description of each layer, including: Determine the domain type of the domain image, and match the description dimension and the description definition based on each description dimension from the type-dimension comparison table; The dimension extraction model is matched from the type-definition-identification comparison table according to the description definition, and the domain image is layer extracted to obtain the extraction layer under the corresponding description dimension, and then the meaning of the extraction layer is analyzed to obtain the layer description.

3. The method for generating and fine-tuning multimodal data of domain images according to claim 2, characterized in that: The extracted layer is analyzed for meaning to obtain a layer description, including: Retrieving the process log of layer extraction of the domain image by the dimension extraction model, and determining the extraction focus coefficient of each position point in the corresponding extraction layer respectively; The extraction focus coefficients are divided according to the extraction thresholds matching the corresponding description dimensions, and the surfaces formed by the position points greater than or equal to the extraction threshold are regarded as the first layer, and the surfaces formed by the position points less than the extraction threshold are regarded as the second layer; Performing global feature analysis on the corresponding extraction layer to obtain the first feature, and at the same time, performing global feature analysis on the corresponding first layer and second layer respectively to obtain the corresponding second feature and third feature; Performing natural language processing on the first feature, the second feature, and the third feature respectively to obtain a corresponding first description, a second description, and a third description; Establishing a first difference function between the first description and the second description, a second difference function between the first description and the third description, and a fusion function between the second description and the third description; Obtaining a supplementary feature according to the first difference function, the second difference function and the fusion function; The first feature and the supplementary feature are analyzed for their meanings to obtain corresponding layer descriptions.

4. The method for generating and fine-tuning multimodal data of domain images according to claim 2, characterized in that: Combining the initial description to generate a comprehensive description includes: If there is only one description dimension under the domain type, the corresponding layer description is regarded as the description to be analyzed; If there are multiple description dimensions under the domain type, at this time, the layer descriptions under all description dimensions are fused to obtain the description to be analyzed; A comprehensive description is generated based on the description to be analyzed and combined with the initial description.

5. The method for generating and fine-tuning multimodal data of domain images according to claim 4, characterized in that: The layer descriptions under all description dimensions are fused to obtain the description to be analyzed, including: Extract the quantity words and category words in each layer description to construct a description vector; Based on deep learning technology, the interest definitions of the category words under each description dimension are mined, and the interesting words under the corresponding category words are obtained, and then the expanded vector is constructed; Determine the type of interest for the word of interest under each category of words in the expanded vector, and count the number of occurrences of the same type of interest in the expanded vector; Filter the interesting type with the largest number of statistics and match the corresponding attention mechanism from the type-mechanism comparison table; Based on the attention mechanism under each description layer, the layer descriptions of all extracted layers are embedded with descriptions of interest, and then the descriptions are fused to obtain the description to be analyzed.

6. The method for generating and fine-tuning multimodal data of domain images according to claim 5, characterized in that: Get the words of interest under the corresponding category vocabulary, including: The category words under the description dimension are used as indexes respectively, and the strong related definitions based on each index under the corresponding description dimension are mined from the specified database by combining deep learning technology; A cluster analysis is performed on all strongly related definitions under each index to obtain a definition cluster, and the reference definition of the corresponding definition cluster is regarded as an interest definition, wherein the number of the definition cluster is at least one.

7. The method for generating and fine-tuning multimodal data of domain images according to claim 1, characterized in that: Fine-tuning the visual multimodal model using the multimodal data includes: Converting the multimodal data into a current input sequence that matches a visual multimodal model; Converting a general description of the domain image into an original input sequence that matches the visual multimodal model; Compare and analyze the original input sequence and the current input sequence of the same domain image, determine the difference sequence and the first number Nd of the difference sequence, and then The second number obtained constructs the partition window, where Indicates the rounding up symbol; Indicates from the range Randomly select a positive integer from; Locking the first position of the first occurrence of the difference sequence and the last position of the last occurrence of the difference sequence in the current input sequence to obtain a position sequence segment; If the segment window of the position sequence segment is the same as the size of the partition window, the current input sequence is kept unchanged; If they are inconsistent, the position sequence segments are sequentially divided according to the division window at the first position to obtain a number of combined sequences, and each combined sequence, the sequence before the first position, and the sequence after the last position are respectively combined to obtain a new sequence; Performing natural language processing on the new sequence, if the semantics of the new sequence is similar to the semantics of the current input sequence, retaining the corresponding new sequence; Otherwise, the position where the difference sequence appears for the second time is screened, and the new sequence obtained subsequently is re-judged in combination with the last position where the difference sequence appears for the last time; The visual multimodal model is trained based on all retained new sequences and all current input sequences.

8. The method for generating and fine-tuning multimodal data of domain images according to claim 7, characterized in that: The image description and the converted text are semantically integrated and expanded based on the large language model to generate an initial description, including: Based on the large language model, we search for related concepts in the image description and the converted text and establish corresponding relationships. Based on the corresponding relationship, complementary information in the image description and the converted text is integrated; Using external knowledge graphs, find new concepts and knowledge related to the fused semantics, and combine with the background knowledge of the domain image to determine the depth and breadth of the semantics, and generate integrated and extended texts; The integrated and extended texts are restated in natural language to obtain an initial description.

Citation Information

Patent Citations

  • A multi-scale visual attention image description method

    CN109670576A

  • Scene graph generation method based on multi-modal contrast learning

    CN117710780A

  • Multi-modal metaphor detection method based on hierarchical consistency reasoning and graph description

    CN118643901A

  • Image description method, device and equipment and computer readable storage medium

    CN119339378A

  • Multimodal semantic analysis and image retrieval

    US20240354336A1

Cited By

  • Intelligent data query method and system based on big data

    CN120429479A

  • Intelligent data query method and system based on big data

    CN120429479B

  • Method for analyzing abnormal category of tightening curve based on multi-modal large model

    CN120726407A