Target generalization identification method and device based on virtual data generation
By designing target-style text prompt templates and large visual language models, high-fidelity and diverse virtual data are generated. The mapping network is used to optimize the target content features, achieving generalized target recognition and improving the adaptability and recognition accuracy of deep learning models in real scenarios.
Patent Information
- Application Number
- CN202511263802.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing virtual data generation technologies are difficult to generate high-fidelity and diverse virtual data, and cannot effectively supplement real surgical scene data, resulting in insufficient recognition capabilities of deep learning models in lesion areas.
Design a target-style text prompt template, combine it with the text-image model to generate virtual data, use the visual language model and mapping network to extract target content features, and achieve target generalization recognition through cosine similarity calculation.
Generating high-fidelity and diverse virtual data improves the adaptability and recognition accuracy of the model in real scenarios and enhances the value of virtual data in practical applications.
Smart Images

Figure CN120747490A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a method and device for identifying target generalization based on virtual data generation. Background Art
[0002] In certain scenarios, acquiring real data is extremely difficult, such as capturing images in extreme environments, recording rare events, or in areas involving privacy protection (such as healthcare and security). In these scenarios, real data is often scarce, costly, or even unavailable, severely limiting the training and application of deep learning models. To compensate for the scarcity of real data, virtual data generation techniques are needed, using computer technology to generate large amounts of realistic virtual data. However, traditional virtual data generation typically relies on game engines, simulating real environments through 3D scene modeling and saving images of the 3D scene to generate virtual data. However, this approach is not only complex and costly, and fails to capture the diversity of real scenarios, but also generates data with a style that deviates significantly from the natural distribution of real data.
[0003] In the medical field, early detection and accurate diagnosis are key to improving cure and survival rates for a variety of diseases. However, surgical scene data is difficult to obtain under real-world conditions. Existing surgical scene datasets make it difficult to fully tap the potential of deep learning models and effectively identify lesion areas. Therefore, it is necessary to generate virtual data under surgical scenes to supplement real surgical scene data and fully tap the potential of deep learning models by increasing the amount of data. However, existing virtual data generation technologies usually focus on generating virtual data by simulating real scenes through game engines or 3D modeling. However, the generation process is complex and difficult to adjust dynamically. It cannot quickly adapt to the diverse needs of real scenes, and it is difficult to effectively supplement real surgical scene data, and it cannot significantly improve the model's ability to identify lesion areas.
[0004] Therefore, how to design a target generalization recognition method that can make up for the lack of real data and has high recognition efficiency is a technical problem to be solved. Summary of the Invention
[0005] Based on this, it is necessary to provide a target generalization identification method and device based on virtual data generation to address the problems of the existing technology.
[0006] In a first aspect, an embodiment of the present application discloses a method for identifying target generalization based on virtual data generation, comprising the following steps: S1: Generate prompt text based on the relevant features of the training data of the target in the actual application scenario; S2: Inputting the prompt text into a pre-trained large model of text and graph to generate virtual data of the target in the actual application scenario; S3: Inputting the virtual data into a visual language model to generate visual global features, target pseudo-words, and style pseudo-words of the virtual data; the visual language model includes a CLIP visual encoder, a CLIP text encoder, a target mapping network, a style mapping network, and a content mapping network; S4: embedding the target pseudo-words and style pseudo-words into a predefined target-style text template and embedding the target pseudo-words into a predefined target text template and processing them through a CLIP text encoder respectively to obtain target-style text features and target text features; S5: constructing the visual global feature and the target-style text feature into a first positive sample pair, and enabling the visual language model to learn a target mapping network and a style mapping network based on the first positive sample pair to update the CLIP visual encoder; S6: Inputting the global visual feature projection into the content mapping network to obtain target content features; S7: constructing the target content feature and the target text feature into a second positive sample pair, and causing the content mapping network to perform comparative learning based on the second positive sample pair to update the content mapping network; S8: Calculating target text features of the reference target image based on the updated visual language model, and calculating target content features of the real target image based on the updated visual language model; S9: Calculate the cosine similarity between the target content feature of the real target image and the target text feature of the reference target image, and obtain a target generalization recognition result based on the cosine similarity ranking result.
[0007] Preferably, step S2 includes: S21: Design an image partitioning diagram to specify the generation area of each image data, which is used to store data of the same target in different styles; S22: Acquire the morphological data of the target; S23: Generate multi-style and multi-object training data based on prompt text, image separation box diagram and morphological data; S24: Preprocess the multi-style and multi-object training data and crop the data based on the image partition frame to obtain virtual data.
[0008] Preferably, step S3 includes: S31: Processing the virtual data through the CLIP visual encoder of the visual language large model to obtain visual global features; S32: Projecting the visual global features through a target mapping network to obtain a target pseudo-word; S33: using the output result of the penultimate layer of the CLIP visual encoder as a visual local feature; S34: Calculating the cross attention of the local visual features and the global visual features, and screening the local visual features that are highly relevant to the image semantics according to the attention scores; S35: Project the filtered visual local features through the style mapping network to obtain style pseudo-words.
[0009] Preferably, step S4 includes: S41: Design two pseudo-word-based text templates, namely the target-style text template and the target text template; S42: Processing the target-style text template and the target text template respectively by the CLIP text encoder to obtain target-style text template features and target text template features; S43: embedding the target pseudo-words and the style pseudo-words into target-style text template features to obtain target-style text features; S44: Embed the target pseudo-word into the target text template feature to obtain the target text feature.
[0010] In a second aspect, an embodiment of the present application discloses a device for identifying target generalization based on virtual data generation, comprising: A prompt text generation unit, configured to generate a prompt text based on relevant features of the training data of the target in the actual application scenario; A virtual data generating unit, configured to input the prompt text into a pre-trained large model of text and graphs to generate virtual data of the target in the actual application scenario; a virtual processing unit, configured to input the virtual data into a visual language macro model to generate visual global features, target pseudo-words, and style pseudo-words of the virtual data; the visual language macro model includes a CLIP visual encoder, a CLIP text encoder, a target mapping network, a style mapping network, and a content mapping network; a text feature generation unit, configured to embed the target pseudo-words and the style pseudo-words into a predefined target-style text template and embed the target pseudo-words into a predefined target text template and process them respectively through a CLIP text encoder to obtain target-style text features and target text features; a visual encoder updating unit, configured to construct a first positive sample pair from the visual global feature and the target-style text feature, and enable the visual language model to learn a target mapping network and a style mapping network based on the first positive sample pair to update the CLIP visual encoder; a content feature generating unit, configured to input the global visual feature projection into the content mapping network to obtain target content features; a content mapping network updating unit, configured to construct the target content feature and the target text feature into a second positive sample pair, and enable the content mapping network to perform comparative learning based on the second positive sample pair to update the content mapping network; an image feature calculation unit, configured to calculate target text features of a reference target image based on the updated visual language large model, and to calculate target content features of a real target image based on the updated visual language large model; The target generalization recognition unit is used to calculate the cosine similarity between the target content feature of the real target image and the target text feature of the reference target image, and obtain the target generalization recognition result based on the cosine similarity sorting result.
[0011] Compared with the existing technology, the present invention has the following beneficial effects: based on the target generalization recognition method generated by virtual data, a target-style text prompt template is designed and combined with a large model of text and image to generate high-fidelity and diverse virtual data; using the large model of visual language and the mapping network, the target content features are extracted and optimized, and the influence of style features on the target content features is reduced; by calculating the cosine similarity between the target content features of the real target image and the target text features of the reference target image, the target generalization recognition of the real target image is achieved, the adaptability and recognition accuracy of the model trained with virtual data in real scenes are improved, and strong support is provided for the value mining of virtual data in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] A more complete understanding of the exemplary embodiments of the present invention can be obtained by referring to the following drawings. The drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present invention and do not constitute a limitation of the present invention. In the drawings, the same reference numerals generally represent the same components or steps.
[0013] Figure 1 A flowchart of a method for identifying target generalization based on virtual data generation provided in an embodiment of the present application; Figure 2 A schematic diagram of a target generalization identification device based on virtual data generation provided by an embodiment of the present application; DETAILED DESCRIPTION
[0014] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0015] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0016] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0017] Reference Figure 1 This embodiment discloses a method for identifying target generalization based on virtual data generation, comprising the following steps: S1: Generate prompt text based on the relevant features of the training data of the target in the actual application scenario; Specifically, by summarizing the types of targets that need to be identified in actual application scenarios and the types of styles that these targets may exhibit due to the influence of other factors, a dedicated text prompt template is designed to construct prompt text for generating multi-style, multi-target virtual data. The prompt text content includes fixed descriptive words, feature descriptive words, and decorative descriptive words. Among them, fixed descriptive words are used to control the quality of virtual data generation and describe common features in actual application scenarios, making the generated virtual data closer to the real scene; feature descriptive words are used to control the representative characteristics of each target, making the generated virtual data more biased towards the specified target; decorative descriptive words are used to expand the fixed descriptive words and feature descriptive words, so that the generated virtual data exhibits different styles and increases diversity. Through different combinations of fixed descriptive words, feature descriptive words, and decorative descriptive words, a variety of prompt texts for generating multi-style, multi-target virtual data are constructed.
[0018] S2: Input the prompt text into the pre-trained text-based graph model to generate virtual data of the target in the actual application scenario; Specifically, the prompt text obtained in step S1 is combined with the image segmentation boxes (the image segmentation boxes are pre-designed, resulting in an image similar to a 5×2 table. The edge detection model in ControlNet and the Lora multi-view generation model are used to control the large model, generating each image within each segmentation box) and target morphology data (the target morphology data is derived from an existing dataset. The pose detection model in ControlNet extracts information about the possible poses of the target in the existing dataset to generate more realistic and natural virtual data). This guides the large model to generate virtual data that is both realistic and easy to compare and observe. The image segmentation boxes are arranged in 2 rows and 6 columns. The Canny edge detection model is used to identify the segmentation boxes and restrict the model to generating virtual data within the boxes. The image segmentation boxes consist of a black border and a white inner image. The edge detection algorithm distinguishes the position of the boxes and controls the generation of virtual data according to the order of the boxes.
[0019] Furthermore, the OpenPose model is used to identify the target's morphological data from the provided reference target data, guiding the model to generate the morphological data for virtual data. The resulting prompt text is fed into the Stable Diffusion text-based model, which, combined with the image separation box diagram and the target morphological data, imposes constraints to generate virtual data of different targets in different styles. The generated virtual data is manually screened to remove data with inadequate image generation, poor quality, or that does not meet the requirements of practical application scenarios. The screened images are cropped according to the image separation box diagram to generate a virtual dataset suitable for training.
[0020] S3: Input virtual data into the visual language model to generate the visual global features, target pseudo-words, and style pseudo-words of the virtual data; the visual language model includes the CLIP visual encoder, CLIP text encoder, target mapping network, style mapping network, and content mapping network; Specifically, target pseudo-words, style pseudo-words, and content pseudo-words are essentially three types of features that represent different meanings. Because they are features rather than specific texts, they are called "pseudo-words." In CLIP (the visual language model of this application), after calculation by CLIP's visual encoder and text encoder, the image and text are unified into a feature space. Therefore, the features originally from the image can also be regarded as text features in this space. The specific meaning of these three pseudo-words is: using visual features to enhance text representation, so that text features can just express the information contained in the image, avoiding the disadvantage of simply using text to fully express visual information. The significance of using a mapping network is that the original image features contain too much redundant information, and the mapping network is used to streamline the information related to specific content. Because the essence of pseudo-words is features, when embedding a template, the template is first calculated into features and then embedded, and then the text features are calculated again after embedding, rather than directly embedding the pseudo-words into the template and then calculating the text features.
[0021] Specifically, the process of using the CLIP visual language model to calculate image and text features and the visual encoder of the CLIP model to calculate the visual global features of virtual data can be expressed as: (1); in, Represents the visual global features of virtual data, The visual encoder representing the CLIP visual language large model, Represents the input virtual data. The target mapping network is used to learn the target pseudo-word features. The process of projecting the visual global features into the target pseudo-word using the target mapping network can be expressed as: (2); in, represents the target pseudoword, Represents the target mapping network. The process of calculating the visual local features of virtual data using the visual encoder of the CLIP visual language model can be expressed as: (3); in, Represents the visual local features of virtual data, represents the visual encoder of the visual language model excluding the last layer, and L represents the number of layers of the visual encoder.
[0022] Next, we calculate the cross-attention scores of the local and global visual features of the virtual data and select the top k local visual features that are highly relevant to the image semantics. This process can be expressed as: (4); (5); (6); in, , , represents the learnable weight matrix, represents the query vector, represents the key vector, represents the value vector, A represents the attention weight matrix, represents the attention weight after filtering, represents the corresponding value vector, Represents the filtered visual local features.
[0023] A multi-layer perceptron (MLP) is used as a style mapping network to learn style pseudo-word features. The process of using the style mapping network to project visual local features into style pseudo-words can be expressed as: (7); in, Pseudo-words expressing style, Represents the style mapping network.
[0024] S4: embedding the target pseudo-words and style pseudo-words into a predefined target-style text template and embedding the target pseudo-words into a predefined target text template and processing them through the CLIP text encoder respectively to obtain target-style text features and target text features; Specifically, the pseudo-word based target-style text template adopts the following format: “A [x] style of aphoto of [y].”, while the pseudo-word based target text template adopts the following format: “A photo of [y].”.
[0025] Among them, the square brackets "[]" indicate that pseudo words need to be embedded, x represents the style pseudo word, and y represents the target pseudo word. The process of calculating the text template features can be expressed as: (8); (9); in, The text encoder representing the CLIP visual language model, Represents the target-style text template features, represents a target-style text template, Represents the target text template features, Represents the target text template.
[0026] Use the target pseudowords obtained in step S3 and style pseudowords Embedding target-style text template features The process can be expressed as: (10); in, Represents target-style text features. It is a concatenation function used to embed pseudo-words into templates. The calculation process is to connect the pseudo-word features and template features together to form a feature sum. Embed target text template features The process can be expressed as: (11); in, Represents the target text features.
[0027] S5: The global visual features and the target-style text features are constructed as the first positive sample pair, so that the visual language model learns the target mapping network and the style mapping network based on the first positive sample pair to update the CLIP visual encoder; Specifically, fine-tuning the visual encoder, object mapping network, and style mapping network of the CLIP visual language large model is performed in two steps.
[0028] By analyzing the visual global features of the virtual data obtained in step S3 and target-style text features Perform contrastive learning, fine-tune the visual encoder of the large visual language model, learn the target-style mapping network and the style mapping network, and use the contrastive loss: (12); in, Represents each sample in a batch of data, all subscripts with Refers to the first image in a batch of images. Images, refers to calculating the contrast loss of each image, and then adding the contrast loss of all images together to get the contrast loss of this batch of images, and then backpropagating the optimization model. Similarly, the subscript in the denominator Represents the calculation of the When calculating the similarity of an image and all text features one by one, the calculation is done up to the text features; is the similarity function, is the temperature parameter; N is the batch size, specifically how much data is computed at once. The computation is performed batch by batch, not one image at a time. This is why the contrastive loss maximizes the similarity between each image and its corresponding text features within the batch, while minimizing the similarity between each image and its corresponding texts within the batch.
[0029] The contrastive learning process uses a contrastive loss function to calculate the similarity between global visual features and target-style text features. Minimizing this loss function aims to make the global visual features of the image being calculated as similar as possible to its corresponding target-style text features, while also minimizing the similarity to the target-style text features of other images. This process optimizes the model by minimizing the loss function through backpropagation. Because of the use of a contrastive loss function, this step is called contrastive learning. The loss function measures the difference between the current model output and the true result, and the fine-tuning process is the process of minimizing the loss function. Fine-tuning all model parameters is performed automatically through backpropagation.
[0030] S6: Project the global visual features into the content mapping network to obtain the target content features; Specifically, a multi-layer perceptron (MLP) is used as a content mapping network to learn target content features. The process of using the content mapping network to project the visual global features obtained in step 3 into the target content feature space can be expressed as: (13); in, Indicates the target content characteristics, Represents a content mapping network.
[0031] S7: constructing the target content feature and the target text feature into a second positive sample pair, and causing the content mapping network to perform comparative learning based on the second positive sample pair to update the content mapping network; Specifically, by analyzing the target content features of virtual data and target text features Perform contrastive learning to learn the content mapping network so that the content mapping network can effectively learn the target content features. The contrast loss used is: (14); S8: Calculating target text features of the reference target image based on the updated visual language model, and calculating target content features of the real target image based on the updated visual language model; S9: Calculate the cosine similarity between the target content features of the real target image and the target text features of the reference target image, and obtain the target generalization recognition result based on the cosine similarity ranking result.
[0032] Specifically, the target text features of the reference target image are calculated by the visual encoder and target-style mapping network of the visual language model trained in step S5. , the target content features of the real target image are calculated through the visual encoder of the visual language model trained in step S5 and the content mapping network trained in step S7 , calculate the target content features of the real target image and the target text features of the reference target image The cosine similarity between them is calculated and sorted by similarity to obtain the target generalization recognition result.
[0033] Specifically, the reference target image is used in object recognition to define the target to be identified. This concept originates from the task itself and is equivalent to providing the target to be identified. The real target image is the image used in actual applications. Virtual data is generated to compensate for the difficulty in obtaining real data and is used only for model training. When the model is applied in practice, that is, to verify the model's generalization to real scenes, images from real scenes are required, rather than the virtual data used during training. During learning, the model focuses solely on extracting generalizable features. Since the positive examples used in contrastive learning are images and text, a reference target image is not required. However, during inference, the target to be found by the model must be defined. In general contrastive learning, this contrast target is text (category). However, this method does not involve text during inference; text is only used for training. The reference target image is used as the contrast target to find the same target in different images. Directly using images as target definitions avoids the problem of similar targets being difficult to distinguish simply using text, and enables more fine-grained recognition of generalized targets. The reference target image is equivalent to the text used in inference in general contrastive learning. The real target image is the image used in actual applications.
[0034] Specifically, the shape of the content feature of the real target image is [B, D], which means B images, each with D-dimensional features; the shape of the target text feature of the reference target image is [N, D], which means N reference target images / candidates (multi-image recognition, if only one image is recognized, N=1), each of which also has D-dimensional features. After calculating the cosine similarity, the shape is [B, N]. Row represents the The similarity between an image and all N texts; Column represents the The similarity between a text and all B images. If there is one reference target image, the shape of the cosine similarity matrix is [B, 1], that is, each value in the first column represents the similarity between a real target image and the reference target image, and the sorting is performed accordingly. The purpose of this method is to identify images that have the same target as the reference target image from multiple real target images given one (or more) reference target images. The essence of obtaining a generalized recognition result based on cosine similarity sorting is to match the real target image that is most similar to the reference target image, and the final result is a real target image that contains the same target as the reference target image. In other words, there is no need to input text information, only a reference target image is needed, and images containing the same target can be found from other images.
[0035] This application is based on a target generalization recognition method generated by virtual data. It designs a target-style text prompt template and combines it with a large model of text and image to generate high-fidelity and diverse virtual data; it uses a large visual language model and a mapping network to extract and optimize target content features, reducing the influence of style features on target content features; by calculating the cosine similarity between the target content features of the real target image and the target text features of the reference target image, it achieves target generalization recognition of the real target image, improves the adaptability and recognition accuracy of the model trained with virtual data in real scenarios, and provides strong support for the value mining of virtual data in practical applications.
[0036] Based on the same inventive concept, the present application also provides an apparatus for implementing the aforementioned method. The solution provided by the apparatus is similar to the solution described in the aforementioned method. Therefore, the specific limitations of the embodiment of the target generalization identification apparatus based on virtual data generation provided below can be found in the above-mentioned limitations of the target generalization identification method based on virtual data generation, and will not be repeated here.
[0037] Reference Figure 2 This embodiment discloses a target generalization identification device 20 based on virtual data generation, comprising: The prompt text generating unit 201 is used to generate prompt text based on relevant features of the training data of the target in the actual application scenario; A virtual data generating unit 202 is configured to input the prompt text into a pre-trained large model of a text graph to generate virtual data of the target in the actual application scenario; A virtual processing unit 203 is configured to input the virtual data into a visual language model to generate visual global features, target pseudo-words, and style pseudo-words of the virtual data; the visual language model includes a CLIP visual encoder, a CLIP text encoder, a target mapping network, a style mapping network, and a content mapping network; A text feature generating unit 204 is configured to embed the target pseudo-words and the style pseudo-words into a predefined target-style text template and embed the target pseudo-words into a predefined target text template and process them through a CLIP text encoder respectively to obtain target-style text features and target text features; A visual encoder updating unit 205 is configured to construct a first positive sample pair from the visual global feature and the target-style text feature, and enable the visual language model to learn a target mapping network and a style mapping network based on the first positive sample pair to update the CLIP visual encoder; A content feature generating unit 206 is configured to input the global visual feature projection into the content mapping network to obtain target content features; a content mapping network updating unit 207 configured to construct the target content feature and the target text feature into a second positive sample pair, and enable the content mapping network to perform comparative learning based on the second positive sample pair to update the content mapping network; An image feature calculation unit 208 is configured to calculate target text features of a reference target image based on the updated visual language model, and to calculate target content features of a real target image based on the updated visual language model; The target generalization identification unit 209 is configured to calculate the cosine similarity between the target content feature of the real target image and the target text feature of the reference target image, and obtain a target generalization identification result based on the cosine similarity ranking result.
[0038] It should be noted that the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0039] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0040] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.
[0041] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0042] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0043] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application, and they should all be included in the scope of the claims and description of the present application.
Claims
1. A target generalization identification method based on virtual data generation, characterized in that: The steps include: S1: Generate prompt text based on the relevant features of the training data of the target in the actual application scenario; S2: Inputting the prompt text into a pre-trained large model of text and graph to generate virtual data of the target in the actual application scenario; S3: Inputting the virtual data into a visual language model to generate visual global features, target pseudo-words, and style pseudo-words of the virtual data; the visual language model includes a CLIP visual encoder, a CLIP text encoder, a target mapping network, a style mapping network, and a content mapping network; S4: embedding the target pseudo-words and style pseudo-words into a predefined target-style text template and embedding the target pseudo-words into a predefined target text template and processing them through a CLIP text encoder respectively to obtain target-style text features and target text features; S5: constructing the visual global feature and the target-style text feature into a first positive sample pair, and enabling the visual language model to learn a target mapping network and a style mapping network based on the first positive sample pair to update the CLIP visual encoder; S6: Inputting the global visual feature projection into the content mapping network to obtain target content features; S7: constructing the target content feature and the target text feature into a second positive sample pair, and causing the content mapping network to perform comparative learning based on the second positive sample pair to update the content mapping network; S8: Calculating target text features of the reference target image based on the updated visual language model, and calculating target content features of the real target image based on the updated visual language model; S9: Calculate the cosine similarity between the target content feature of the real target image and the target text feature of the reference target image, and obtain a target generalization recognition result based on the cosine similarity ranking result.
2. The method according to claim 1, characterized in that Step S2 includes: S21: Design an image partitioning diagram to specify the generation area of each image data, which is used to store data of the same target in different styles; S22: Acquire the morphological data of the target; S23: Generate multi-style and multi-object training data based on prompt text, image separation box diagram and morphological data; S24: Preprocess the multi-style and multi-object training data and crop the data based on the image partition frame to obtain virtual data.
3. The method according to claim 2, characterized in that Step S3 includes: S31: Processing the virtual data through the CLIP visual encoder of the visual language large model to obtain visual global features; S32: Projecting the visual global features through a target mapping network to obtain a target pseudo-word; S33: using the output result of the penultimate layer of the CLIP visual encoder as a visual local feature; S34: Calculating the cross attention of the local visual features and the global visual features, and screening the local visual features that are highly relevant to the image semantics according to the attention scores; S35: Project the filtered visual local features through the style mapping network to obtain style pseudo-words.
4. The method according to claim 3, characterized in that Step S4 includes: S41: Design two pseudo-word-based text templates, namely the target-style text template and the target text template; S42: Processing the target-style text template and the target text template respectively by the CLIP text encoder to obtain target-style text template features and target text template features; S43: embedding the target pseudo-words and the style pseudo-words into target-style text template features to obtain target-style text features; S44: Embed the target pseudo-word into the target text template feature to obtain the target text feature.
5. A target generalization recognition device based on virtual data generation, characterized in that: include: A prompt text generation unit, configured to generate a prompt text based on relevant features of the training data of the target in the actual application scenario; A virtual data generating unit, configured to input the prompt text into a pre-trained large model of text and graphs to generate virtual data of the target in the actual application scenario; a virtual data processing unit, configured to input the virtual data into a visual language macro model to generate visual global features, target pseudo-words, and style pseudo-words of the virtual data; the visual language macro model includes a CLIP visual encoder, a CLIP text encoder, a target mapping network, a style mapping network, and a content mapping network; a text feature generation unit, configured to embed the target pseudo-words and the style pseudo-words into a predefined target-style text template and embed the target pseudo-words into a predefined target text template and process them respectively through a CLIP text encoder to obtain target-style text features and target text features; a visual encoder updating unit, configured to construct a first positive sample pair from the visual global feature and the target-style text feature, and enable the visual language model to learn a target mapping network and a style mapping network based on the first positive sample pair to update the CLIP visual encoder; a content feature generating unit, configured to input the global visual feature projection into the content mapping network to obtain target content features; a content mapping network updating unit, configured to construct the target content feature and the target text feature into a second positive sample pair, and enable the content mapping network to perform comparative learning based on the second positive sample pair to update the content mapping network; an image feature calculation unit, configured to calculate target text features of a reference target image based on the updated visual language large model, and to calculate target content features of a real target image based on the updated visual language large model; The target generalization recognition unit is used to calculate the cosine similarity between the target content feature of the real target image and the target text feature of the reference target image, and obtain the target generalization recognition result based on the cosine similarity sorting result.
Citation Information
Patent Citations
Image generation method and device, medium and computer program product
CN118505858A
Method for diversifying style words through entropy maximization to realize domain generalization
CN118587723A
Method, computer device, and non-transitory computer-readable recording medium for generating image
US20240371049A1
Text-guided multi-modal relationship extraction method and apparatus
WO2025130069A1