Long-tail data generation method and computer equipment

By combining detail extraction, entity recognition, and causal relationship analysis with a multimodal large language model and diffusion model, the method addresses the issues of realism and diversity in long-tail traffic scene image generation in existing technologies, thereby improving the understanding and generalization capabilities of autonomous driving models in complex scenarios.

CN120953403APending Publication Date: 2025-11-14SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510862239.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies struggle to generate realistic and diverse long-tail traffic scene images, resulting in insufficient generalization ability of autonomous driving models in complex scenarios. Furthermore, existing methods are unable to effectively understand the various relationships and semantics within complex scenarios.

Method used

Long-tail data is generated through detail extraction, entity recognition, abnormal entity acquisition, and causal relationship analysis. A combination of multimodal large language models and diffusion models is used to generate the data, including multi-round counterfactual analogies and multi-model generation and feedback mechanisms, to ensure the authenticity and diversity of the images.

Benefits of technology

It improves the understanding and generalization ability of autonomous driving models in long-tail scenarios, generates realistic and diverse images, meets the data requirements of autonomous driving for complex scenarios, and provides an effective evaluation benchmark for long-tail scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953403A_ABST
    Figure CN120953403A_ABST
Patent Text Reader

Abstract

The invention relates to a long-tail data generation method and computer equipment, which focus on acquisition of abnormal entities and capture of association between the abnormal entities and other entities, so that a formed benchmark can comprehensively and accurately measure the understanding ability of a model to a long-tail scene when evaluating the long-tail driving scene. And an effective long-tail scene evaluation benchmark can be formed. A non-single text generation image is obtained by replacing an entity and adding a new entity, important semantic information significance is well understood, a complex scene is truly generated, and key information is grasped. Compared with the prior art, the method for generating the long-tail data and the computer equipment disclosed by the invention have the advantages that the purposes of enabling the generated image to be real in a complex scene and reducing the homogenization of the generated image can be realized. In addition, the understanding ability of rare and key long-tail scenes can be improved, the blank of scarcity of long-tail scene data is filled through a data synthesis mode, and the adaptability and safety in complex and rare scenes are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis and generation technology, and in particular to a method and computer device for generating long-tail data. Background Technology

[0002] With the rise and development of multimodal large-scale models, their powerful reasoning capabilities are considered a new direction for achieving intelligence in the field of autonomous driving. Compared to traditional black-box driving systems, large-scale models have better logical thinking capabilities and interpretability. However, to achieve autonomous driving in real-world scenarios, the long-tail problem is often encountered. These long-tail scenarios are rare but crucial to the safety of autonomous driving. Furthermore, due to the scarcity of data, such examples are rarely collected in the training set, resulting in poor generalization ability of the model in these scenarios. These problems limit the development of autonomous driving technology in complex real-world scenarios and urgently need to be addressed.

[0003] Data generation is considered a potential solution to this problem, especially with the increasing controllability and quality of diffusion models. However, existing technologies face numerous challenges in generating images of complex traffic scenes. First, the generated images are often unrealistic. For example, when using text-to-image models (such as DALL-E-3) to generate an image from a realistic description like "a chair fell from a truck," the simple description often leads to incorrect spatial relationships, and the chair's position may not reflect reality. Conversely, if MLLMs are used to generate a detailed description before generating the image, the excessive detail can cause the model to focus on unimportant details while neglecting long-tail semantic relationships, thus failing to generate valuable long-tail images.

[0004] Secondly, the generated images are highly homogenized. Existing methods mostly mimic the same query image, resulting in a lack of diversity in the generated images. The text corpus is also limited. Using these homogenized images to train MLLMs will limit the model's generalization ability in open driving scenarios and may even lead to overfitting.

[0005] The root of these problems lies in the difficulty of existing models in handling the diverse relationships and semantics within complex scenes. Solving these problems presents challenges such as enabling models to effectively understand complex spatial relationships, avoiding the generation of irrelevant objects, and increasing the diversity of generated images. Furthermore, acquiring the rich, high-level semantics required to guide the generation of long-tail traffic images also presents a challenge, making it difficult for existing technologies to meet the demands of autonomous driving for generating long-tail traffic scene data. Summary of the Invention

[0006] To address the challenges of enabling models to effectively understand complex spatial relationships, avoid generating irrelevant objects, and increase the diversity of generated images, as well as the need to acquire rich high-level semantics to guide the generation of long-tail traffic images, which makes existing technologies insufficient to meet the requirements of autonomous driving for generating long-tail traffic scene data, this invention proposes a method for generating long-tail data.

[0007] The technical solution adopted in this invention is a method for generating long-tail data, comprising:

[0008] Detail extraction: Extracting text details from long-tail scene information;

[0009] Entity recognition, identifying individual entities and related information within long-tail scene information;

[0010] Abnormal entity acquisition: Abnormal entities are acquired based on relevant and prior information about the entity.

[0011] Causal relationship analysis: analyzing the causal relationships between anomalous entities and other entities;

[0012] Long-tail data generation involves generating text descriptions based on causal relationships, and then generating long-tail data based on these text descriptions.

[0013] In some embodiments, the entity recognition step includes:

[0014] Obtain normal scene information under normal circumstances based on the category of long-tail scene information;

[0015] Identify each entity in both normal and long-tail scene information.

[0016] In some embodiments, the entity recognition step includes:

[0017] Identify relevant prompts based on long-tail scenario information;

[0018] The multimodal large language model is guided by prompt words to identify various entities in normal scene information and long-tail scene information;

[0019] Furthermore, a multimodal large language model is used to classify the relevant information of entities based on their different characteristics.

[0020] In some embodiments, the long-tail data generation step includes:

[0021] After generating the text description, the text description is processed to obtain refined text that focuses on key events and relationships, and long-tail data is generated based on the refined text.

[0022] In some embodiments, the long-tail data generation step includes:

[0023] Propose counterfactual hypotheses for anomalous entities in the text description, replace the entities in the text description and / or generate new entities in the text description, and obtain the updated text;

[0024] Generate long-tail data based on the updated text.

[0025] In some embodiments, after the step of obtaining the updated text, the method includes:

[0026] Annotate the consequences of updating the text to obtain initial analysis;

[0027] By combining the initial analysis and historical synthesized text information, iterative text is generated after iteration;

[0028] Generate long-tail data based on iterated text.

[0029] In some embodiments, the step of generating long-tail data based on the updated text includes:

[0030] Set the long-tail data as image data;

[0031] Use at least two generative models to generate target images with different biases for the updated text.

[0032] In some embodiments, after the step of generating target images with different biases for updating text using at least two generative models, the method includes:

[0033] The target image is scored based on the corresponding patterns and common sense of long-tail events to obtain an evaluation score;

[0034] Filter out target images whose evaluation scores are below a threshold;

[0035] Unfiltered target images are selected as preferred target images for long-tail events.

[0036] In some embodiments, after filtering target images with evaluation scores below a threshold, the method includes:

[0037] For target images below a threshold, analyze the reasons for failure and obtain feedback parameters;

[0038] The updated text is modified based on feedback parameters.

[0039] To address the shortcomings of existing computer devices in handling multiple relationships and semantics in complex scenarios, this invention proposes a computer device.

[0040] The technical solution adopted by the present invention is a computer device, comprising: a processor and a memory, wherein the memory is used to store computer program code, the computer program code including computer instructions, and the computer device executes the method described above when the processor executes the computer instructions.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] This application discloses a method for generating long-tail data, including steps such as detail extraction, entity recognition, abnormal entity acquisition, causal relationship analysis, and long-tail data generation. It focuses on acquiring abnormal entities and capturing the relationships between abnormal entities and other entities, enabling the resulting benchmark to comprehensively and accurately measure the model's understanding of long-tail driving scenarios. This allows existing evaluation metrics and datasets to fully consider the specificities of long-tail scenarios, such as the localization of rare events and risk assessment. This allows researchers to accurately judge the model's performance in long-tail scenarios, providing targeted guidance for model improvement and optimization. Simultaneously, by replacing entities and adding new entities, it obtains non-single text-generated images, better understanding important semantic information, making the generation of complex scenarios realistic, and capturing key information. Furthermore, most current generation methods rely on imitating the same query image, resulting in highly similar images in content and structure. This application, however, introduces a diverse driving mechanism in the generation process, using different generation models to allow the generated images to vary across a wide range, thereby meeting the diverse data requirements of autonomous driving scenarios. Non-homogeneous images generated from corpora, when used to train MLLMs, do not cause the model to overfit to specific scene patterns, thereby improving its generalization ability in open driving scenarios to cope with various unknown long-tail scenarios.

[0043] Compared with the prior art, the method for generating long-tail data disclosed in this application can form an effective evaluation benchmark for long-tail scenes, make the generated images realistic in complex scenes, and reduce the homogenization of generated images.

[0044] This application discloses a computer device that can improve the computer device's understanding of rare but critical long-tail scenarios. By synthesizing data, it fills the gap of scarce data in long-tail scenarios and enhances the adaptability and security of the computer device in complex and rare scenarios. Attached Figure Description

[0045] The present invention will now be described in detail with reference to the embodiments and accompanying drawings, wherein:

[0046] Figure 1 A flowchart illustrating a method for generating long-tail data according to an embodiment of the present invention is shown;

[0047] Figure 2 A flowchart illustrating a method for generating long-tail data in long-tail events in the transportation sector is shown.

[0048] Figure 3 It shows according to Figure 2 A more detailed flowchart illustrating a method for generating long-tail data in long-tail events within the transportation sector is provided.

[0049] Figure 4 A flowchart illustrating a method for generating long-tail data when performing image analysis and analog feedback in long-tail events in the transportation field is shown.

[0050] Figure 5 It shows according to Figure 4 A more detailed flowchart illustrating a method for generating long-tail data when performing image analysis and analog feedback in long-tail events in the transportation sector is provided.

[0051] Figure 6 The image shows long-tailed event images output using multiple generative models;

[0052] Figure 7 The diagram shows a long-tailed data plot generated by validating multiple multimodal large language models.

[0053] Figure 8 A comparison chart showing the number of images generated with and without an analogy process is provided. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Examples of embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar components or components having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0055] How to address the issue of homogeneous data generation and improve the diversity of synthetic data? This problem has never been explored before. Gradually, we began to consider letting the model make counterfactual assumptions about objects in long-tail scenes, such as: what if it wasn't a chair that fell, but a box? Would the consequences be the same? This allows us to explore various possibilities or other similar scenarios. This step relies on the powerful thinking ability of large models, much like human brainstorming, which is something that current single-function neural networks cannot achieve. This application proposes a long-tail data generation method specifically for exploring various possibilities or other similar scenarios.

[0056] The technical solution adopted in this invention is as follows, please refer to... Figure 1 A method for generating long-tail data, comprising:

[0057] S10, Detail Extraction: Extract text details of long-tail scene information;

[0058] S20, Entity recognition: Recognizing each entity and related information of the entity in long-tail scene information;

[0059] S30. Abnormal entity acquisition: Obtain abnormal entities based on relevant and prior information of the entities.

[0060] S40. Causal relationship analysis: Analyze the causal relationships between anomalous entities and other entities;

[0061] S50, Long-tail data generation: Generate text descriptions based on causal relationships, and generate long-tail data based on the text descriptions.

[0062] The new evaluation benchmark, DriveTail, is specifically designed for evaluating long-tail driving scenarios. It incorporates rich long-tail traffic scenario data and features detailed image annotations and risk scoring. Compared to existing benchmarks, DriveTail can more comprehensively and accurately assess a model's understanding of long-tail driving scenarios, providing strong support for improving model performance.

[0063] Long-tail scene information can be presented as images, text, characters, codes, audio, or other types of electronic data. The process involves first extracting textual details from the input long-tail scene information, and then identifying various entities from it. When the long-tail scene information is an image, this method, compared to directly capturing abnormal regions in the long-tail scene image, provides a more comprehensive view of all the information in the image. It also lays the foundation for establishing connections between abnormal entities and other entities. Furthermore, after textual detail extraction, the generation of text descriptions and long-tail data is not affected by the original long-tail scene image, making it easier to generate unique images and enhance diversity. When the long-tail scene information is text, textual detail extraction is even more convenient. It should be noted that when the long-tail scene information is characters or codes, textual detail extraction requires prior decoding, compilation, or pre-training on a relevant training set based on the type of character or code. This is particularly important for acquiring long-tail data that is not directly obtained in real-world applications but is processed. Furthermore, when the long-tail scene information is in the form of characters or codes, the entities subsequently obtained are also described through text, and the entities obtained may correspond to a part of the pattern in the characters or a sequence in the codes.

[0064] Furthermore, in long-tail events, there must be an actor who initiates the action, as well as the attributes and characteristics of each subject. By extracting text details and converting them into text form, we can more clearly obtain the various entities in the long-tail scene information, and quickly identify abnormal and unreasonable parts of the scene information. This is more effective than directly identifying a single subject in the long-tail scene information.

[0065] It should be noted that this application searches for anomalous entities. Compared to directly searching for anomalous events, anomalous entities are the basic units that constitute anomalous events. Anomalous events revolve around anomalous entities. Therefore, searching for anomalous entities can most appropriately decompose anomalous events and obtain the most accurate textual descriptions of long-tail events.

[0066] In addition to identifying anomalous entities, it is also necessary to analyze the causal relationships between anomalous entities and other entities. This is to elaborate and describe in detail the reasons for the occurrence of anomalous entities and the impact they will have on other entities. This is especially important for obtaining accurate text descriptions. The more detailed the description of causal relationships, the more accurate the text description will be.

[0067] This application does not limit the type of long-tail data generated. Preferably, the type of long-tail data can be the same as the type of long-tail scene information. For example, when the input long-tail scene information is an image, the generated long-tail data will also be an image. The type of long-tail data can also be images, text, characters, codes, or other types of electronic data.

[0068] In some particularly specific embodiments, please refer to Figure 2 and Figure 3 A method for generating long-tail data includes image observation and image extraction, wherein image observation includes image detail extraction and traffic entity classification, and image extraction includes causal relationship analysis and regenerated description.

[0069] Image detail extraction (O-Step 1): Long-tail traffic scenes are complex, containing many key details and rare concepts. To gain a deeper understanding of the images, this study uses GPT-4o to perform a detailed description of each query image X∈DM, forming T. O1 Taking the scenario of "a chair falling off a truck" as an example, GPT-4o will provide detailed information about the truck's color, model, and driving status, the chair's appearance and location, as well as the surrounding environment, such as weather conditions, road type, and the location of other vehicles, providing a solid foundation for subsequent steps.

[0070] Traffic Entity Classification (O-Step 2): To clearly present the query scenario structure, this study uses specific prompts to guide MLLM to classify T. O1 And the actual situation T GTThe basic traffic participants in the scene are categorized and organized. These entities cover environmental and road conditions, vehicles, people, and other objects. Each entity can potentially affect driving behavior. By listing them in categories, the elements and their attributes in the scene, such as location, color, and size, can be clearly displayed, facilitating subsequent analysis of the interactions between entities.

[0071] Causal Analysis (Step 1): In this stage, the anomalous elements in the given scenario need to be identified, and causal analysis needs to be performed. Leveraging the powerful reasoning capabilities of MLLM and combining prior traffic knowledge, the study identifies the most anomalous entity from the entity list as the anomalous entity, such as "the chair fell from the truck." MLLM further analyzes its causal relationships with other participating entities. For example, because the chair was not properly secured to the truck, it fell while the vehicle was moving, posing a potential danger to other vehicles on the road, thus providing a detailed description of the anomalous event.

[0072] Regenerating the Description (I-Step 2): The analyzed text obtained in I-Step 1 is too complex and unsuitable for image generation by the diffusion model. Therefore, this study uses Prompt to allow GPT-4o to modify the long text and generate a concise new description T. I2 And highlight the key long-tail event E k For example, a fallen chair. In this way, the generator can focus on key events and relationships, enabling the generated images to better mimic the long-tail semantic causal relationships of the query images.

[0073] For example, when the basic fact is that a chair has fallen from a truck ahead, the following steps can be taken:

[0074] Step 1: Use GPT4o to describe the long-tailed image in detail and obtain relevant descriptive text. Examples include a scene on a highway, a white pickup truck carrying a large piece of furniture that is not properly secured, the fall location being on the left, a silver sedan, trees and leaves, other vehicles, and sunny weather.

[0075] Step 2: Use GPT4o again to organize these descriptions into entities with concise characteristics and list them by category. For example, the Environment and Roads category includes entities such as highway and sunny day; the Vehicles category includes entities such as white pickup truck, silver sedan, and other vehicles; the People category has no entities; and the Other Objects category includes entities such as an outdoor chair, tree, and elevated highway sign.

[0076] Step 1: Input these entities into GPT4o to determine what is unusual in the image within the traffic scene and analyze the reasons in detail. The anomaly analysis output at this point might be: The unusual part in the image is a chair improperly secured to a white pickup truck. This causes it to lift off the ground while the vehicle is moving. This creates a sense of instability and potential danger for other drivers on the road.

[0077] Step 2: Input the analyzed reasons back into GPT4o. Focusing on the reasons for the anomalies analyzed above, summarize the above text information into a shorter text message. The output short text message could be: A white pickup truck is driving on the highway, carrying a large, unsecured aerial chair. On a sunny day, a silver sedan drives beside it.

[0078] The above brief text information is input into the DALL-E-3 model to generate the image;

[0079] Step A-Step 1 involves generating images using multiple rounds of counterfactual analogies to brief textual information, replacing or generating entities using GPT4o without altering the semantic meaning of the image. Possible outputs include: First-round analogy: A motorcycle in the right lane with a large ladder horizontally placed at its rear, the ladder's ends dangerously extending into the adjacent lane; Second-round analogy: A red truck speeding along a country road with a large suitcase flying through the air, sliding out from its spread-out rear; Third-round analogy: A blue truck in the middle lane carrying a large mattress, the mattress billowing in the wind, dangerously secured to the roof with several straps.

[0080] Step A-2 involves using multiple image generation models to generate images. During the generation process, GPT4o is used to identify the authenticity and reliability of the generated images and obtain feedback information. Based on the feedback information, GPT4o reprocesses the generated text information and imports it back into the original generation model to generate images again, in order to obtain the optimal result.

[0081] In some embodiments, the entity recognition step includes:

[0082] Obtain normal scene information under normal circumstances based on the category of long-tail scene information;

[0083] Identify each entity in both normal and long-tail scene information.

[0084] In order to obtain entities that exist in most common situations and avoid situations where these common entities do not appear in certain long-tail scene information, the entity recognition step not only identifies entities in long-tail scene information but also identifies entities in normal scene information.

[0085] In some embodiments, the entity recognition step includes:

[0086] Identify relevant prompts based on long-tail scenario information;

[0087] The multimodal large language model is guided by prompt words to identify various entities in normal scene information and long-tail scene information;

[0088] Furthermore, a multimodal large language model is used to classify the relevant information of entities based on their different characteristics.

[0089] Among them, the multimodal large language model can receive text, images, voice, video and even sensor data. It can form a good entity recognition capability regardless of whether the type of information in the long-tail scene is certain or uncertain. It can also be applied to the detail extraction step for text detail extraction.

[0090] To maximize the capabilities of the multimodal large language model, this application determines prompt words based on long-tail scene information, thereby obtaining a high relevance to the long-tail scene information and better predicting long-tail events that may occur in the scene, and matching appropriate keywords.

[0091] Furthermore, classifying entity-related information based on different entities using a multimodal large language model can also improve the efficiency of subsequent abnormal entity acquisition steps.

[0092] In some embodiments, the long-tail data generation step includes:

[0093] After generating the text description, the text description is processed to obtain refined text that focuses on key events and relationships, and long-tail data is generated based on the refined text.

[0094] This is because when using MLLMs to generate detailed descriptions to highlight long-tail events, excessive detail makes it difficult for the generative model to focus on key long-tail relationships, resulting in images containing incorrect objects or illogical scenes. This is because the model cannot effectively filter and integrate key semantics when processing large amounts of information, causing the generated results to deviate from reality.

[0095] In some embodiments, the long-tail data generation step includes:

[0096] Propose counterfactual hypotheses for anomalous entities in the text description, replace the entities in the text description and / or generate new entities in the text description, and obtain the updated text;

[0097] Generate long-tail data based on the updated text.

[0098] In a particularly specific embodiment, please refer to Figure 4 and Figure 5 Image generation methods based on long-tail events include multi-round counterfactual analogy and multi-model generation and feedback.

[0099] Multi-turn Counterfactual Analogy (A-Step 1): To increase the diversity of generated images and improve the fine-tuning performance of downstream models, this study investigates analogical thinking through multi-turn dialogue with GPT-4o. In each iteration, based on the logical analysis of the observation and imitation phases, counterfactual reasoning is used to explore more analogical scenarios. This is applied to the core entity E. k Counterfactual hypotheses are proposed, allowing the model to replace or generate new entities, such as replacing a chair with a suitcase, while preserving long-tail semantic relationships. The model also considers and labels the consequences of these changes. In subsequent iterations, the model combines initial image analysis with historical synthesized text information to explore more possibilities and generate new scene descriptions. The initial images are the original images used for analysis; this data is collected from real-world scenes and is relatively small in quantity. The synthesized text information is a new scene description generated by GPT through multiple rounds of analogy analysis of the long-tail semantics of the images.

[0100] Multi-model generation and feedback (A-Step 2): To further improve image diversity and quality, a multi-model generation and feedback mechanism was adopted. For a single synthesized text, three high-performance generation models—DALL-E-3, FLUX, and MUSES—were selected to generate images, each with its own characteristics and capable of producing images with varying degrees of bias. An independent GPT-4o was used as a discriminator to evaluate and score the generated images based on the physical laws and common sense of traffic scenes. If an image score was below a threshold θ, it was filtered, and the feedback reason was sent to the GPT-4o agent in A-Step 1, prompting it to modify the analogy text and regenerate the image, thus ensuring the realism and reliability of the final generated image. It should be noted that the threshold θ was selected using GPT as the discriminator. The prompt function allowed the model to judge whether the scene in the generated image conformed to common sense and to score the scene's realism, relying on the large model's own experience and world knowledge to complete the evaluation. This threshold was selected empirically to balance the speed and quality of image generation.

[0101] When using diffusion models to generate images of complex scenes, there is a gap between the complexity of the text description and the model's processing capabilities. Text-to-image models, such as DALL-E-3, tend to generate unrealistic images, such as incorrect chair placement, when given a simple description like "a chair fell off a truck." This is because they lack a deep understanding of the complex spatial relationships and physical logic within the scene. This is to address the contradiction between text description and image generation.

[0102] Innovative Generative Framework: DriveCoach, a novel training-free data synthesis framework, is proposed to generate long-tail traffic scene images based on the "Observation, Imitation, Analogy" (OIA) thought chain (CoT). Unlike existing methods that directly utilize diffusion models or simply rely on MLLMs to generate images, DriveCoach effectively integrates the advantages of MLLMs and diffusion models through a structured process, progressively generating realistic and diverse long-tail traffic images.

[0103] A unique generation process: In the observation phase, MLLM is used to perform a top-down detailed deconstruction of long-tail traffic images to obtain comprehensive scene information; in the imitation phase, MLLM is used to deeply analyze the causal relationships of abnormal events, highlighting long-tail relationships and generating targeted descriptions; in the analogy phase, the reasoning ability of MLLM is used to conduct multiple rounds of counterfactual analogies, combined with multi-model generation and discriminative feedback mechanisms, to increase image diversity and realism. This unique three-step generation process differs from the single generation mode in existing technologies and is more in line with human cognition and creative logic for complex scenes.

[0104] The method disclosed in this application can be used not only for long-tail generation in traffic scenarios, but also by using large models as agents to analyze various complex application scenarios, simulate human thinking processes, and generate more long-tail data. Of course, different generation requirements require different downstream task verification, but they are all extensions of the ideas in this application.

[0105] In some embodiments, after the step of obtaining the updated text, the method includes:

[0106] Annotate the consequences of updating the text to obtain initial analysis;

[0107] By combining the initial analysis and historical synthesized text information, iterative text is generated after iteration;

[0108] Generate long-tail data based on iterated text.

[0109] The consequences of updating text refer to the impact and results on long-tail events when replacing entities or adding new entities. An initial analysis is performed on the impact and results, and based on the initial analysis and the previously generated text information, including the text descriptions, refined text, and other updated text mentioned above, all of which can be used together with the initial analysis to generate iterative text.

[0110] In some embodiments, the step of generating long-tail data based on the updated text includes:

[0111] Set the long-tail data as image data;

[0112] Use at least two generative models to generate target images with different biases for the updated text.

[0113] Specifically, to further improve image diversity and quality, the study employs a multi-model generation and feedback mechanism. For a single synthesized text, three high-performing generation models—DALL-E-3, FLUX, and MUSES—are selected to generate images, each with its own characteristics and capable of producing images with varying degrees of bias.

[0114] In some embodiments, after the step of generating target images with different biases for updating text using at least two generative models, the method includes:

[0115] The target image is scored based on the corresponding patterns and common sense of long-tail events to obtain an evaluation score;

[0116] Filter out target images whose evaluation scores are below a threshold;

[0117] Unfiltered target images are selected as preferred target images for long-tail events.

[0118] This study utilizes an independent GPT-4o discriminator to evaluate and score generated images based on the physical laws and common sense of traffic scenarios. Images with scores below a threshold θ are filtered out. This filtering method eliminates target images with low realism and reliability, thereby ensuring the quality of the selected target images.

[0119] In some embodiments, after filtering target images with evaluation scores below a threshold, the method includes:

[0120] For target images below a threshold, analyze the reasons for failure and obtain feedback parameters;

[0121] The updated text is modified based on feedback parameters.

[0122] The feedback reason is then sent to the GPT-4o agent in A-Step 1, prompting it to modify the analog text and regenerate the image, thereby ensuring the authenticity and reliability of the final generated image.

[0123] In some embodiments, after the multi-model generation step, the following is included:

[0124] The selected generative models are at least two of the DALL-E-3 model, FLUX model, and MUSES model.

[0125] The verification of the image generation method based on long-tailed events is as follows:

[0126] Please refer to Figure 6 After obtaining targeted long-tail text descriptions through our observation and analysis, we applied our multiple generative models to generate long-tail event images, which are then compared with some other previous generation methods. Qualitative experiments show that our method more reasonably reflects the semantics of long-tail events while also achieving realistic generation.

[0127] For evaluation on the long-tail dataset DriveTail, we selected some of the most popular and best-performing multimodal large models for validation. Please refer to [link / reference]. Figure 7 First, we compare zero-shot data, then we gradually add a small amount of real long-tail data, and finally we use synthetic long-tail data generated by different generation methods.

[0128] As can be seen, with the help of the data generated by our method, some smaller models, such as LLaVa-Next and Shikra, outperformed large models like GPT-4o.

[0129] Furthermore, compared to data generated directly using previous methods, our method achieved better results in downstream task feedback.

[0130] Please refer to Figure 8 Compared to homogeneous generation without an analogy process, our method not only achieves higher results when we increase the number of long-tail event images generated, but also avoids overfitting caused by too many duplicate image samples.

[0131] To address the shortcomings of existing computer devices in handling multiple relationships and semantics in complex scenarios, this invention proposes a computer device.

[0132] The technical solution adopted in this invention is a computer device, comprising: a processor and a memory, wherein the memory is used to store computer program code, the computer program code including computer instructions, and the computer device executes the above-described method when the processor executes the computer instructions.

[0133] The computer device includes a processor and memory. Optionally, the computer device also includes an input device and an output device. The processor, memory, input device, and output device are coupled together via connectors, which include various interfaces, transmission lines, or buses, etc., and are not limited in this application embodiment. It should be understood that in the various embodiments of this application, coupling refers to mutual connection in a specific way, including direct connection or indirect connection through other devices, such as through various interfaces, transmission lines, buses, etc.

[0134] The processor may include one or more processors, such as one or more central processing units (CPUs). If the processor is a CPU, it can be a single-core CPU or a multi-core CPU. Optionally, the processor may be a processor group consisting of multiple CPUs, with the multiple processors coupled to each other via one or more buses. Optionally, the processor may also be other types of processors, etc., which are not limited in the embodiments of this application.

[0135] The memory can be used to store computer program instructions, as well as various types of computer program code, including program code for executing the scheme of this application. Optionally, the memory includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), which is used for related instructions and data.

[0136] Input devices are used to input data and / or signals, and output devices are used to output data and / or signals. Input and output devices can be independent devices or an integrated device.

[0137] It is understood that in this embodiment of the application, the memory can be used not only to store related instructions, but also to store related data. This embodiment of the application does not limit the specific data stored in the memory.

[0138] In practical applications, computer devices may also include other necessary components, including but not limited to any number of input / output devices, processors, memory, etc., and all computer devices that can implement the embodiments of this application are within the protection scope of this application.

[0139] In the description of this specification, the use of terms such as "Embodiment 1," "this embodiment," or "in one embodiment" indicates that the specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example; moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in one or more embodiments or examples.

[0140] In the description of this specification, the terms "connection," "installation," "fixing," "setting," and "having" are interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0141] In the description of this specification, relational terms such as “first” and “second” are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0142] The above description of the embodiments is intended to enable those skilled in the art to understand and apply the technology of this invention. Those skilled in the art can easily make various modifications to these examples and apply the general principles described herein to other embodiments without creative effort. Therefore, this invention is not limited to the above embodiments. Modifications in the following situations should be within the scope of protection of this invention: ① New technical solutions implemented based on the technical solution of this invention and combined with existing common knowledge, where the technical effects of the new technical solution do not exceed the technical effects of this invention; ② Equivalent substitutions of some features of the technical solution of this invention using known technology, resulting in the same technical effects as those of this invention; ③ Extendable technical solutions based on the technical solution of this invention, where the substantive content of the extended technical solution does not exceed the technical solution of this invention; ④ Equivalent transformations made using the content of this specification and drawings, directly or indirectly applied to other related technical fields.

Claims

1. A method for generating long-tail data, characterized in that, include: Detail extraction: Extracting text details from long-tail scene information; Entity recognition, identifying each entity in the long-tail scene information and related information of the entity; Abnormal entity acquisition: Abnormal entities are acquired based on the relevant information and prior information of the entity. Causal relationship analysis, analyzing the causal relationships between the abnormal entity and other entities; Long-tail data generation involves generating text descriptions based on the causal relationships, and then generating long-tail data based on these text descriptions.

2. The method for generating long-tail data according to claim 1, characterized in that, The steps in entity recognition include: Obtain normal scene information under normal circumstances based on the category of the long-tail scene information; Identify each entity in the normal scene information and the long-tail scene information.

3. The method for generating long-tail data according to claim 2, characterized in that, The steps in entity recognition include: Based on the long-tail scene information, relevant prompt words are determined; The prompt words guide the multimodal large language model to identify each entity in the normal scene information and the long-tail scene information; The multimodal large language model is used to classify the relevant information of the entities according to their differences.

4. The method for generating long-tail data according to claim 1, characterized in that, The steps involved in generating long-tail data include: After generating the text description, the text description is processed to obtain refined text that focuses on key events and relationships, and the long-tail data is generated based on the refined text.

5. A method for generating long-tail data according to any one of claims 1 to 4, characterized in that, The steps involved in generating long-tail data include: Counterfactual hypotheses are proposed for the anomalous entities in the text description, and the entities in the text description are replaced and / or new entities are generated in the text description to obtain the updated text. The long-tail data is generated based on the updated text.

6. The method for generating long-tail data according to claim 5, characterized in that, After obtaining the updated text, the steps include: The consequences of the updated text are labeled to obtain initial analysis; By combining the initial analysis and the historically synthesized text information, the iterative text is generated. The long-tail data is generated based on the iterative text.

7. The method for generating long-tail data according to claim 5, characterized in that, The step of generating the long-tail data based on the updated text includes: Set the long-tail data as image data; At least two generative models are used to generate target images with different biases for the updated text.

8. The method for generating long-tail data according to claim 7, characterized in that, After the step of generating target images with different biases for the updated text using at least two generative models, the process includes: The target image is scored based on the corresponding patterns and common sense of long-tail events to obtain an evaluation score; Filter out target images whose evaluation scores are below a threshold; The unfiltered target image is selected as the preferred target image for long-tail events.

9. A method for generating long-tail data according to claim 8, characterized in that, After filtering the target images whose evaluation scores are below a threshold, the process includes: For the target images that are below the threshold, perform failure reason analysis and obtain feedback parameters; The updated text is modified using the feedback parameters.

10. A computer device, characterized in that, include: A processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein, when the processor executes the computer instructions, the computer device performs the method as described in any one of claims 1 to 9.