Data generation method and device based on dynamic multi-level cue word, equipment and medium

By preprocessing multimodal input data and dynamically adjusting hierarchical prompt words, the generative model can more accurately generate high-quality multimodal scene data required in the field of embodied intelligence, solving the problems of insufficient data consistency and authenticity in existing technologies, and achieving the effect of flexible adaptation to different generation needs.

CN120596718APending Publication Date: 2025-09-05PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510764795.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The quality of data generated by existing generative models in the field of embodied intelligence depends on the diversity and quantity of training data. They lack data consistency and authenticity, lack dynamic adjustment mechanisms, and are unable to flexibly generate specific types of data, resulting in low accuracy of generated data.

Method used

By receiving multimodal input data and performing preprocessing, preprocessed data is generated; initial prompt words are received and graded and adjusted to generate multi-level prompt words that are adapted to the generation model; and the generation model is used to perform data generation processing based on the preprocessed data and the multi-level prompt words to generate target scene data.

Benefits of technology

It improves the accuracy of data generation, solves the problems of data scarcity, poor consistency in multimodal generation, and insufficient data in extreme scenarios in the field of embodied intelligence, achieves efficient generation of high-quality multimodal scenario data, improves the consistency and authenticity of generated data, and flexibly adapts to different generation needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596718A_ABST
    Figure CN120596718A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to the fields of financial science and technology and medical health, and discloses a data generation method, device, equipment and medium based on dynamic multi-level cue words, which are applied to financial market dynamic simulation and risk prediction scenes and multi-modal medical image auxiliary diagnosis generation scenes. The method comprises the steps that multi-modal input data are received, the multi-modal input data are preprocessed, preprocessed data are generated, and the multi-modal input data comprise environment sensing data and action sequence data; receiving an initial cue word, and grading and adjusting the initial cue word to generate a multi-level cue word adaptive to the generation model; and performing data generation processing based on the preprocessed data and the multi-level cue words through a generation model to generate target scene data. According to the method, the problems of scarcity of data in the intelligent field, poor multi-modal generation consistency and insufficient extreme scene data are solved, and the accuracy of data generation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data generation method, device, equipment and medium based on dynamic multi-level prompt words. Background Art

[0002] Amid the rapid development of artificial intelligence (AI), embodied intelligence, a key branch of AI, is dedicated to empowering intelligent entities with the ability to perceive, learn, make decisions, and act in physical environments. AI technology allows for the generation of corresponding content using prompts, a practice that has become a vital part of our daily lives and work. This data generation method can be applied to the fields of fintech and healthcare, for example, in the simulation and risk prediction of financial market dynamics, as well as in the generation of multimodal medical image-assisted diagnosis. It can generate financial market simulation content and risk predictions based on user-entered multi-level prompts, or generate corresponding medical images and case studies based on user-entered multi-level prompts related to healthcare.

[0003] The existing data generation method uses a generative model to generate data content based on a single-level prompt word. However, the existing generative model still faces some challenges when applied to the field of embodied intelligence. First, the quality of the data generated by the generative model depends on the diversity and quantity of the training data. If the training data is insufficient or unevenly distributed, the generated data may be biased or distorted. Secondly, when processing multimodal data (such as text, images, videos, etc.), the generative model often finds it difficult to ensure the consistency and authenticity of the data. Especially when simulating complex scenarios, the generated data may not fully reflect the complexity of the real world. In addition, the existing generative model lacks a dynamic adjustment mechanism and cannot flexibly generate specific types of data according to specific needs. As a result, when facing different tasks, the generated data may not meet the needs of model training. Therefore, the consistency and authenticity of the data generated by the existing data generation method are low, resulting in low accuracy of data generation. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a data generation method, device, equipment and medium based on dynamic multi-level prompt words to improve the accuracy of data generation.

[0005] In order to solve the above technical problems, the present invention provides a data generation method based on dynamic multi-level prompt words, including:

[0006] receiving multimodal input data and preprocessing the multimodal input data to generate preprocessed data, wherein the multimodal input data includes environmental perception data and action sequence data;

[0007] receiving an initial prompt word, and grading and adjusting the initial prompt word to generate a multi-level prompt word adapted to the generative model;

[0008] The generation model performs data generation processing based on the pre-processed data and the multi-level prompt words to generate target scene data.

[0009] In order to solve the above technical problems, the embodiment of the present application provides a data generation device based on dynamic multi-level prompt words, including:

[0010] a data preprocessing module, configured to receive multimodal input data and preprocess the multimodal input data to generate preprocessed data, wherein the multimodal input data includes environmental perception data and action sequence data;

[0011] a multi-level prompt word adjustment module, configured to receive initial prompt words, and classify and adjust the initial prompt words to generate multi-level prompt words adapted to the generation model;

[0012] The target scene data generation module is used to perform data generation processing based on the pre-processed data and the multi-level prompt words through the generation model to generate target scene data.

[0013] To solve the above technical problems, the present invention adopts a technical solution: providing a computer device including one or more processors; and a memory for storing one or more programs, so that the one or more processors implement any of the above-mentioned methods for generating data based on dynamic multi-level prompt words.

[0014] To solve the above technical problems, the present invention adopts a technical solution: a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the above-mentioned methods for generating data based on dynamic multi-level prompt words.

[0015] The embodiments of the present invention provide a data generation method, device, equipment, and medium based on dynamic multi-level prompt words. The method includes: receiving multimodal input data and preprocessing the multimodal input data to generate preprocessed data, wherein the multimodal input data includes environmental perception data and action sequence data; receiving initial prompt words and grading and adjusting the initial prompt words to generate multi-level prompt words adapted to a generation model; and performing data generation processing based on the preprocessed data and the multi-level prompt words by the generation model to generate target scene data. The embodiments of the present invention solve the problems of data scarcity, poor multimodal generation consistency, and insufficient extreme scene data in the field of embodied intelligence through multimodal data preprocessing, dynamic graded prompt word adjustment, and multi-level conditional control of the generation model. The embodiments of the present invention have the advantages of efficiently generating high-quality multimodal scene data, improving the consistency and authenticity of generated data, and flexibly adapting to different generation requirements, which is conducive to improving the accuracy of data generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 This is a schematic diagram of an application environment of a data generation method based on dynamic multi-level prompt words in one embodiment of the present invention;

[0018] Figure 2 This is a flowchart of the implementation process of the data generation method based on dynamic multi-level prompt words provided in an embodiment of the present application;

[0019] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S1;

[0020] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step S2;

[0021] Figure 5 yes Figure 4 A schematic flow chart of a specific implementation of step S23;

[0022] Figure 6 yes Figure 2 A schematic flow chart of a specific implementation of step S3;

[0023] Figure 7 yes Figure 6 A schematic flow chart of a specific implementation of step S32;

[0024] Figure 8 This is a flowchart of a data generation method based on dynamic multi-level prompt words provided in another embodiment of the present application;

[0025] Figure 9 Schematic diagram of a data generation device based on dynamic multi-level prompt words provided in an embodiment of the present application;

[0026] Figure 10 It is a schematic diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0028] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0029] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0030] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0031] It should be noted that the data generation method based on dynamic multi-level prompt words provided in the embodiment of the present application is generally executed by a server. Accordingly, the data generation device based on dynamic multi-level prompt words is generally configured in the server.

[0032] The data generation method based on dynamic multi-level prompt words provided by the embodiment of the present invention can be applied in the following aspects: Figure 1In an application environment, a client communicates with a server via a network. The server can receive an initial multi-level prompt word from the client; adapt and generate a target multi-level prompt word based on the initial multi-level prompt word, and generate target scene data based on the multi-level prompt word. The server in the present invention sends the target scene data to the client. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablet computers, and portable wearable devices. The server can be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0033] The data generation method based on dynamic multi-level prompt words provided in the embodiment of the present application can be applied to financial market dynamic simulation and risk prediction scenarios, as well as to multimodal medical image-assisted diagnosis generation scenarios.

[0034] See also Figure 2 , Figure 2 A specific implementation of a data generation method based on dynamic multi-level prompt words is shown.

[0035] It should be noted that the method of the present invention is not limited to the method of Figure 2 The process sequence shown is limited to the following steps:

[0036] S1: Receive multimodal input data and preprocess the multimodal input data to generate preprocessed data, wherein the multimodal input data includes environmental perception data and action sequence data.

[0037] Specifically, multimodal input data is obtained, wherein the multimodal input data includes environmental perception data and action sequence data. Environmental perception data includes image data, text data, point cloud data, etc. Action sequence data includes joint motion trajectory, grasping action data, navigation path data, etc. The multimodal input data is preprocessed to generate preprocessed data. For example, the point cloud data in the multimodal input data is voxelized, and the image data in the multimodal input data is resolution adjusted and normalized. The formatted data is enhanced (such as rotation, cropping, noise addition, etc.) to expand the data set size and obtain preprocessed data.

[0038] See also Figure 3 , Figure 3 A specific implementation of step S1 is shown, which is described in detail as follows:

[0039] S11: Receive the multimodal input data.

[0040] S12: voxelizing the point cloud data in the multimodal input data, adjusting the resolution and normalizing the image data in the multimodal input data, and vectorizing the text data in the multimodal input data to generate initial processed data.

[0041] S13: performing data enhancement on the initial processed data to obtain enhanced data, and performing data expansion on the enhanced data by using a time series interpolation method to generate the pre-processed data.

[0042] Voxelization converts discrete point clouds in three-dimensional space into a regular three-dimensional grid of voxels. This can be achieved using voxel gridding methods supported by three-dimensional convolutional neural networks. By mapping point cloud coordinates to a fixed-size voxel space and calculating the point cloud density within each voxel, this addresses the difficulty in extracting spatial features due to the sparsity of point cloud data. Resolution adjustment scales raw images acquired by different sensors to a uniform pixel size. This can be achieved using a bilinear interpolation algorithm to eliminate the effects of sensor parameter differences on image feature consistency. Normalization maps image pixel values ​​from their original range to a standardized numerical interval. This can be achieved using a maximum-minimum linear transformation method to address variations in image brightness distribution caused by varying lighting conditions. Vectorization converts natural language text into continuous vectors in a semantic space. This can be achieved using the embedding layer of a pre-trained language model. This extracts semantic features from the text to achieve feature space alignment across modal data. Data augmentation perturbs the feature space of the initially processed data. This can be achieved using random rotations, translations, or noise injection to improve the robustness of the generative model to input noise. Temporal interpolation refers to the continuous expansion of discrete sampled data in the time dimension. Specifically, the cubic spline interpolation algorithm can be used to generate intermediate transition data for adjacent frames of the action sequence, thereby enhancing the motion rationality of the generated data in dynamic scenes.

[0043] Specifically, the point cloud in the multimodal input data is converted into a three-dimensional voxel grid through voxelization, retaining the spatial structural features while reducing the data dimension; the image data is processed through resolution unification and normalization to eliminate the impact of sensor differences on visual features; the text data is converted into semantic vectors through vectorization to achieve cross-modal feature alignment. Subsequently, random geometric transformations and noise injection data enhancement operations are applied to the initially processed data to improve the model's adaptability to input perturbations. A temporal interpolation method is further used to generate continuous transition frames between discrete sampling points in the action sequence, extending the temporal continuity of the data. This staged processing framework significantly improves the consistency of multimodal data representation through modality-specific transformation and joint optimization.

[0044] Traditional methods use independent processing flows for multimodal data, which makes it difficult to align the feature spaces of each modality. For example, point cloud data often uses direct downsampling to lose spatial information, image data is simply scaled to ignore lighting differences, and text data uses a bag-of-words model that cannot capture semantic associations. The present application establishes a point cloud spatial structure representation through voxelization, combines image standardization with text semantic embedding to achieve cross-modal feature alignment, and then optimizes data distribution through data enhancement and time interpolation to effectively solve the problem of feature inconsistency caused by modal differences. The embodiment of the present application realizes efficient and unified processing of multimodal input data, eliminates the interference of sensor differences and data sparsity on feature extraction, and enhances the generative model's ability to jointly model spatial structure, visual features, and semantic information in complex scenes. At the same time, the model's anti-noise ability is improved through data enhancement, and temporal interpolation expands temporal continuity, so that the generated scene data is significantly improved in terms of spatial rationality and dynamic coherence.

[0045] S2: receiving initial prompt words, and grading and adjusting the initial prompt words to generate multi-level prompt words adapted to the generation model.

[0046] Specifically, the multi-level prompt word consists of at least three levels, including overall category, scene attributes, and detailed description. An initial prompt word is received and judged to see if it meets the preset level requirements. If so, the initial prompt word is encoded into a recognizable conditional vector for the generative model. If not, the initial prompt word is graded and adjusted to generate a multi-level prompt word that is suitable for the generative model.

[0047] Among them, the overall category refers to the macro-classification of the data generation target, which can be implemented by using category labels described in natural language, such as weather classification or traffic scene classification in autonomous driving scenarios. Its role is to provide global semantic constraints for the generation model. Scene attributes refer to the description of the environmental state or behavioral characteristics of the target scene. Specifically, they can be implemented using structured parameters or semantic vectors, such as light intensity and object movement speed. Its role is to refine the scene feature expression of the generation model. Detailed description refers to the supplementary explanation of specific elements or local features in the scene. Specifically, it can be implemented using a combination of keywords or short sentences, such as vehicle color and pedestrian posture. Its role is to enhance the local authenticity of the generated data.

[0048] See also Figure 4 , Figure 4 A specific implementation of step S2 is shown, which is described in detail as follows:

[0049] S21: receiving the initial prompt word, and determining whether the initial prompt word meets a preset level requirement, and obtaining a determination result.

[0050] S22: If the judgment result is that the initial prompt word meets the preset level requirement, the initial prompt word is encoded into a recognizable condition vector of the generation model to generate the multi-level prompt word.

[0051] S23: If the judgment result is that the initial prompt word does not meet the preset level requirement, the initial prompt word is graded and adjusted to generate the multi-level prompt word adapted to the generation model.

[0052] The preset hierarchical requirements refer to pre-set semantic completeness standards for multi-level prompt words. These can be implemented using a semantic coverage assessment model or rule-matching algorithm to determine whether the initial prompt word has the semantic support to generate the required scene data. Recognizable conditional vectors map text prompt words to numerical representations that can be processed by the generative model. This can be implemented using a word embedding model or an attention encoder, converting semantic information into control parameters for the model's latent space.

[0053] Specifically, when the initial prompt word contains the complete three levels of overall category, scene attributes, and detailed description, it is directly converted into a conditional vector through the semantic encoder, preserving the original semantic information while avoiding redundant processing. If the initial prompt word lacks the scene attribute level, for example, containing only the overall category "rainy day" and the detailed description "ground reflection", the large language model analyzes semantic relevance and supplements the intermediate level scene attribute "low visibility". The weight adjustment module strengthens the influence of this level on the generation process, ultimately forming a complete prompt word structure containing three levels. This process adopts a dynamic hierarchical mechanism, resolving the missing level problem through semantic reconstruction while ensuring generation efficiency. This enables the generative model to adaptively adjust the generation strategy based on the different levels of completeness of the input prompt.

[0054] Traditional methods typically employ a fixed-level prompt word structure, such as using only a single category label or unstructured text descriptions. This results in the generative model being unable to effectively distinguish semantic information of varying granularity. However, the present embodiment significantly improves the expressive power of prompt words for complex scene features by forcing a three-level structure: overall category, scene attributes, and detailed descriptions, combined with a dynamic semantic completion mechanism. For example, when generating extreme traffic accident scenarios, existing techniques can miss key elements in the generated data due to ambiguous prompt word hierarchies. However, by supplementing the scene attribute "vehicle collision angle" and the detailed description "airbag deployment status," the present embodiment makes the generated data more consistent with real-world physical laws. The present embodiment effectively addresses the generation bias caused by insufficient semantic hierarchy in the initial prompt words. Through structured hierarchical division and a dynamic adjustment mechanism, the generative model accurately captures complete semantic information, from macroscopic scene classification to microscopic detailed features, improving the consistency of generated data with the target scene requirements. Furthermore, this method reduces the reliance on manually designed complete prompt words. Even when the input prompt words contain missing information, high-quality scene data can still be generated through semantic completion, enhancing the generalization ability of the generative model for complex tasks.

[0055] See also Figure 5 , Figure 5 A specific implementation of step S23 is shown, which is described in detail as follows:

[0056] S231: If the judgment result is that the initial prompt word does not meet the preset level requirement, the initial prompt word is parsed by a large language model, and the parsed initial prompt word is structurally divided according to the semantic level to obtain a divided prompt word.

[0057] S232: Adjust the weights of the prompt words at each level in the divided prompt words to obtain adjusted prompt words.

[0058] S233: Encode the adjusted prompt word into a recognizable conditional vector of the generation model to generate the multi-level prompt word.

[0059] Large language model parsing involves using a pre-trained model with natural language understanding capabilities to perform semantic decomposition and intent recognition on the initial prompt word. This can be achieved using a Transformer-based model, for example, by using a self-attention mechanism to capture contextual dependencies within the prompt word. Semantic hierarchical structure partitioning involves classifying the parsed semantic units into three levels: overall category, scene attributes, and detailed description. This can be achieved using rule matching or clustering algorithms, such as extracting the main components and modifiers through dependency parsing. Weight adjustment involves dynamically allocating the semantic influence of different levels based on the input requirements of the generative model. This can be achieved using an attention weight matrix or learnable parameters, such as controlling the contribution of features at each level in the conditional vector through a gating mechanism. Recognizable conditional vectors involve converting the textual prompt word into a numerical representation that the generative model can process. This can be achieved using word embeddings or projection matrices, for example, by mapping text features to the latent space of the generative model through a multi-layer perceptron.

[0060] In an embodiment of the present application, when the semantic level of the initial prompt word is incomplete, the input text is firstly subjected to deep semantic analysis using a large language model to identify the core entities, action attributes and environmental elements therein. For example, the input prompt word "rainy road" is parsed into "weather state - rainy day" and "scene subject - road". The parsing results are then classified according to the three preset semantic levels, with "road" classified as a general category and "rainy day" classified as a scene attribute, while supplementing the missing detail description level. Subsequently, based on the training data distribution of the generative model, the weight of the scene attribute level is increased so that the generator pays more attention to the impact of weather conditions on the scene data. Finally, the adjusted multi-level text is converted into a low-dimensional vector through a conditional encoder as the constraint input of the generative model. Traditional methods rely on fixed templates or keyword matching for prompt word processing, and are unable to identify implicit semantic relationships, resulting in deviations between the generated data and the expected scene. For example, the existing technology may directly map "rainy road" to a single label, ignoring the association between weather conditions and road surface reflections and vehicle trajectories. The embodiments of the present application can dynamically construct conditional information that conforms to the cognitive structure of the generative model through semantic hierarchical division and weight adjustment. For example, it can strengthen the dynamically changing weather parameters in the scene attribute layer, so that the generated road surface water effect is reasonably associated with the vehicle braking distance. The embodiments of the present application can effectively improve the adaptability of the generative model to incomplete prompt words, ensuring that the generated data is consistent with the real data distribution in terms of scene element integrity and physical law rationality. For example, in the task of generating autonomous driving data, even if the input prompt word lacks a description of the lighting conditions, the shadow effect at dusk can still be supplemented through semantic analysis, and the weight can be adjusted to match the training requirements of the obstacle detection model.

[0061] S3: performing data generation processing based on the pre-processed data and the multi-level prompt words through the generation model to generate target scene data.

[0062] Specifically, the multi-level prompt words and preprocessed data are concatenated to generate joint input data. This joint input data is then fed into the generative model, where a conditional embedding layer generates a latent space distribution based on the multi-level prompt words and generates scene attribute information. A cross-modal attention mechanism is then used to generate initial scene data based on the scene attribute information and the joint input data, and multimodal alignment is performed. Finally, the discriminator is optimized to generate the target scene data. Multi-level prompt words refer to structured conditional inputs divided according to semantic levels. This can be achieved by using a large language model to parse the initial prompt words and then divide them into three levels: overall category, scene attribute, and detailed description. This provides semantic constraints of varying granularity for the generative model.

[0063] See also Figure 6 , Figure 6 A specific implementation of step S3 is shown, which is described in detail as follows:

[0064] S31: performing data splicing on the multi-level prompt words and the pre-processed data to generate joint input data.

[0065] S32: Input the joint input data into the generative model, generate a latent space distribution of the generative model based on the multi-level prompt words through a conditional embedding layer, and generate scene attribute information of the target scene data based on the latent space distribution.

[0066] S33: A cross-modal attention mechanism is used to generate initial scene data based on the scene attribute information and the joint input data, and the initial scene data is subjected to information multimodal alignment processing to generate aligned scene data.

[0067] S34: The target scene data is generated by performing discriminant optimization based on the aligned scene data through the discriminator of the generation model.

[0068] Specifically, the conditional embedding layer is a neural network module that maps text conditional vectors to the latent space of the generative model. This can be achieved by fusing the vectors of each level of multi-level prompt words with the feature maps of preprocessed data using a multi-layer perceptron, which is used to dynamically adjust the generation direction. The cross-modal attention mechanism is an attention calculation layer that establishes dynamic associations between data of different modalities. This can be achieved by using a multi-head attention mechanism to calculate the interaction weights between visual features and text features, which is used to coordinate the collaborative generation of multimodal information. Multimodal alignment processing is a constraint mechanism that eliminates semantic conflicts between different modalities. This can be achieved by using a contrastive learning loss function to force the alignment of the distribution of point cloud, image, and text features in the latent space, which is used to ensure the cross-modal consistency of the generated data. During the generation process, the joint input formed by splicing multi-level prompt words and preprocessed data contains both structured semantic conditions and raw data features, providing a fusion constraint foundation for subsequent generation. The conditional embedding layer encodes multi-level prompt words into a latent space distribution. By adjusting the weight ratio of prompt words at different levels, the generative model is controlled to refine scene attribute features under the overall category constraint. When generating initial scene data, the cross-modal attention mechanism dynamically assigns attention weights to features in different modalities based on the text description. For example, when generating a rainy night road scene, it assigns higher weights to raindrop texture features in the image data and road surface reflectivity features in the point cloud data. Multimodal alignment uses a contrastive loss to ensure consistency between the spatial position and semantic description of the generated data in different modalities. For example, this ensures that the position of the generated 3D vehicle model matches the traffic flow direction described in the text. During adversarial training, the discriminator not only evaluates the authenticity of the generated data but also verifies its semantic alignment with multiple levels of cues. For example, it determines whether the generated image satisfies both the overall category of "urban road" and the scene attributes of "dusk time."

[0069] The embodiment of the present application dynamically adjusts the hierarchical weights of multi-level prompt words, so that the generation model can flexibly control the granularity of scene generation according to task requirements. For example, when training an emergency braking model, the weight of the "sudden obstacle" detail level is enhanced to generate refined scene data including the sudden intrusion of pedestrians. The embodiment of the present application solves the semantic inconsistency problem caused by modal fragmentation in the process of multimodal data generation. Through the cross-modal attention mechanism, a dynamic association of modalities such as vision, text, and point cloud is established to ensure that the generated 3D scene data fully matches the lighting conditions and spatial positions of objects described in the text. At the same time, the dynamic controllability of the generation process is achieved. By adjusting the hierarchical weights of multi-level prompt words, target scene data suitable for different training tasks can be generated on demand. For example, when training an extreme weather perception model, multimodal road scenes under high-intensity rainfall conditions are generated by enhancing the "heavy rain" weight of the scene attribute level.

[0070] In a specific embodiment, the prompt words provided by the user are divided into three levels: (1) overall category, such as: robot navigation scene; (2) scene attributes, such as indoor environment, dynamic obstacles, extreme weather, etc.; (3) detailed description, such as office layout, rainstorm weather details, etc. After selecting the content of the prompt words at each level, the system will call the generative model to generate corresponding data based on these prompt words, for example: generate an image of the office layout, including obstacles such as tables and chairs; generate the navigation path data of the robot in rainstorm weather, including the movement trajectory of pedestrians and raindrop effects.

[0071] See also Figure 7 , Figure 7 A specific implementation of step S32 is shown, which is described in detail as follows:

[0072] S321: Input the joint input data into the generation model, and perform feature fusion based on the identifiable conditional vector of the multi-level prompt word and the feature map of the preprocessed data through the conditional embedding layer to adjust the spatial distribution.

[0073] S322: In a generative adversarial network, the identifiable conditional vector is used as an input channel of a generator so that the generator generates initial scene attribute information.

[0074] S323: The initial scene attribute information is gradually denoised through the cross-attention layer of the diffusion model to generate the initial scene attribute information.

[0075] Specifically, a recognizable conditional vector refers to a vectorized conditional representation formed by encoding multi-level prompt words. This can be implemented using word embeddings or Transformer encoders, which serve as semantic constraints to guide the generation process. The generator of a generative adversarial network refers to a deep neural network module used to generate initial scene attribute information. This can be implemented using a convolutional neural network or U-Net structure, where the input channels fuse conditional vectors to enhance semantic relevance. The cross-attention layer of the diffusion model refers to a noise prediction module based on an attention mechanism. This can be implemented using a multi-head cross-attention mechanism. This layer enhances the relevance of multimodal attributes through iterative denoising.

[0076] Specifically, the recognizable conditional vectors of the multi-level cue words are fused with the feature maps of the preprocessed data in the conditional embedding layer. Feature alignment is used to match the semantic-level cues with the multimodal physical features, eliminating the latent space distribution shift caused by modal differences. A generative adversarial network uses the conditional vectors as input channels, forcing the generator to generate initial attribute information in the latent space that is highly relevant to the semantics of the multi-level cues, such as generating road scene textures that conform to specific weather conditions. The diffusion model uses a cross-attention mechanism in the denoising phase to gradually optimize the noise interference in the initial attribute information. For example, through iterative updates, the correlation between vehicle position and road signs is enhanced, ultimately forming an attribute representation that conforms to the real-world scene distribution.

[0077] Existing methods usually adopt a single prompt word or a fixed weight fusion mechanism, which makes it difficult to dynamically adjust the latent space distribution to adapt to the complex relationship of multimodal data. The existing generative adversarial network lacks direct coupling between the conditional vector and the input channel, resulting in insufficient semantic relevance between attribute generation and the prompt word. Existing diffusion models mostly adopt a single-modal denoising strategy, which cannot effectively maintain the consistency of cross-modal attributes. The embodiment of the present application realizes the adaptive adjustment of the latent space distribution of the generative model, solves the problem of inaccurate attribute information and insufficient consistency in the process of multimodal data generation, improves the authenticity and cross-modal matching ability of the generated data in complex scenarios, and provides high-quality multimodal data support for the training of embodied intelligent models.

[0078] See also Figure 8 , Figure 8 A specific implementation after step S3 is shown, which is described in detail as follows:

[0079] S3A: Performing authenticity evaluation on the target scene data through a pre-trained discriminator network, and performing consistency evaluation on the target scene data through a cross-modal matching model to obtain an evaluation result.

[0080] S3B: Adjust the model parameters of the generation model or the level weights of the multi-level prompt words according to the evaluation result.

[0081] The pre-trained discriminator network refers to a neural network classifier optimized through adversarial training. Specifically, it can be implemented using a discriminator based on a convolutional neural network architecture. Its purpose is to identify areas in the generated results that do not conform to physical laws or exhibit mode collapse by comparing the distribution of generated data with real data. The cross-modal matching model refers to a deep learning model capable of multimodal feature alignment. Specifically, it can be implemented by combining a cross-modal attention mechanism with contrastive learning. Its purpose is to verify semantic consistency between different modalities in the generated data, such as whether the position of an object in an image corresponds to the 3D coordinates of the point cloud data. Adjusting the parameters of the generative model based on evaluation results involves updating the weights of the generator through a backpropagation algorithm. Specifically, this can be achieved using gradient descent combined with an evaluation loss function. Its purpose is to optimize the model's latent space representation to improve the generation quality of specific scenes. Adjusting the weights of multi-level prompt words dynamically adjusts the contribution of conditional vectors at different semantic levels. Specifically, this can be achieved using a learnable attention weight allocation mechanism. Its purpose is to balance the generation results between the macroscopic scene layout and the microscopic details.

[0082] Specifically, after the generative model outputs target scene data, a pre-trained discriminator network quantitatively scores the generated data's physical properties, such as texture detail and lighting plausibility. For example, it detects whether the direction of a vehicle's shadow aligns with the position of a light source in an autonomous driving scenario. Simultaneously, the cross-modal matching model cross-validates logical associations between data from different modalities, such as whether the lane pattern in a road image matches the curvature characteristics of the LiDAR point cloud. When a spatial offset is detected between the position of a pedestrian in the image and the obstacle detection results in the point cloud, the cross-modal matching model generates a consistency loss signal. The evaluation results are used to optimize both the generative model parameters and adjust the cue word weights. For model parameter adjustment, the system calculates the weighted sum of the discriminator output and the cross-modal matching loss, and updates the parameters of each generator layer through chained derivation. For cue word weight adjustment, the system dynamically adjusts the weights of the conditional vectors at the three levels—overall category, scene attribute, and detailed description—based on the evaluation results. For example, if an unreasonable scene layout is detected, the weight of the overall category level is increased to strengthen the global constraint.

[0083] Traditional generative model optimization usually relies only on the adversarial training of a single discriminator, and lacks constraints on the inherent correlation of multimodal data. For example, existing technologies may optimize image quality alone while ignoring its correspondence with text descriptions. However, the embodiment of the present application introduces a cross-modal matching model to construct a dual verification mechanism covering the physical properties and semantic logic of the data, solving the common problems of image-text inconsistency and spatial dislocation in multimodal generation scenarios. Existing methods often use fixed-weight prompt words for conditional control, while the dynamic hierarchical weight adjustment mechanism proposed in the embodiment of the present application realizes the coordinated optimization of macro constraints and micro details in the generation process. The embodiment of the present application effectively improves the physical rationality and multimodal consistency of the generated data. For example, in the generation of autonomous driving scenarios, it ensures that the simulated rain and fog weather images match the visibility parameters of the radar point cloud, while maintaining the vehicle motion trajectory in accordance with the laws of dynamics. The embodiment of the application achieves continuous optimization of generation quality through a closed-loop feedback mechanism, providing high-fidelity, strongly correlated multimodal data support for embodied intelligent model training.

[0084] In a specific embodiment, the multi-level prompt words input by the user are (1) overall category: benign or malignant lung nodule discrimination; (2) scene attributes: right upper lobe, 20-year smoking history, elevated CEA tumor marker; (3) detailed description: spicule sign, pleural traction, diameter 8mm; the system generates the following structured report based on the content of the multi-level prompt words: 1. Lesion location: posterior segment of right upper lobe (three-dimensional coordinates x: 120, y: 45, z: 22); 2. Malignancy probability: 82% (based on the spicule sign weight of 40% + smoking history of 30%); 3. Key imaging features: spicule sign length: 2.3mm (visualized after enhancement), pleural indentation depth: 1.8mm (dynamic simulation of traction process); 4. Recommended next examination: PET-CT metabolic value prediction, liquid biopsy.

[0085] In another specific embodiment, the multi-level prompt words entered by the user are (1) Overall category: SME supply chain finance risk rating; (2) Scenario attributes: Manufacturing, annual revenue of 500-1 billion, account period of 90 days; (3) Detailed description: Upstream and downstream concentration > 60%, operating cash flow volatility ± 30%. The risk analysis report output by the system is: 1. Core risk points: Dependence on key customers: The largest customer accounts for 58% (the average for the automotive parts industry is 35%), Cash flow break probability: 23% (if Q3 payment is delayed for more than 15 days); 2. Recommended credit plan: Maximum limit: ¥80 million (receivables must be pledged), Interest rate floating range: LPR + 180bps to + 250bps; 3. Stress test results: Commodity prices rise by 10% → Default probability rises to 31%, Exchange rate breaks 7.2 → Overseas order profit margin compresses to 2.3%.

[0086] In an embodiment of the present application, multimodal input data is received and preprocessed to generate preprocessed data, wherein the multimodal input data includes environmental perception data and action sequence data; initial prompt words are received and graded and adjusted to generate multi-level prompt words adapted to a generative model; and data generation processing is performed by the generative model based on the preprocessed data and the multi-level prompt words to generate target scene data. This embodiment of the present invention solves the problems of data scarcity, poor consistency in multimodal generation, and insufficient data for extreme scenarios in the field of embodied intelligence through multimodal data preprocessing, dynamic graded prompt word adjustment, and multi-level conditional control of the generative model. It has the advantages of efficiently generating high-quality multimodal scene data, improving the consistency and authenticity of generated data, and flexibly adapting to different generation requirements, which is conducive to improving the accuracy of data generation.

[0087] Please refer to Figure 9 , as a response to the above Figure 2 The present application provides an embodiment of a data generation device based on dynamic multi-level prompt words. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0088] like Figure 9 As shown, the data generation device based on dynamic multi-level prompt words of this embodiment includes: a data pre-processing module 41, a multi-level prompt word adjustment module 42 and a target scene data generation module 43, wherein:

[0089] A data preprocessing module 41 is configured to receive multimodal input data and preprocess the multimodal input data to generate preprocessed data, wherein the multimodal input data includes environmental perception data and action sequence data;

[0090] A multi-level prompt word adjustment module 42 is used to receive the initial prompt word, and to classify and adjust the initial prompt word to generate a multi-level prompt word adapted to the generation model;

[0091] The target scene data generating module 43 is configured to generate target scene data by performing data generation processing based on the pre-processed data and the multi-level prompt words through the generation model.

[0092] Furthermore, the data preprocessing module 41 includes:

[0093] a data receiving unit, configured to receive the multimodal input data;

[0094] an initial processing data generating unit, configured to voxelize the point cloud data in the multimodal input data, perform resolution adjustment and normalization processing on the image data in the multimodal input data, and vectorize the text data in the multimodal input data to generate initial processing data;

[0095] The data enhancement unit is used to perform data enhancement on the initial processed data to obtain enhanced data, and to perform data expansion on the enhanced data by using a time series interpolation method to generate the preprocessed data.

[0096] Furthermore, the multi-level prompt words have at least three levels, including overall categories, scene attributes, and detailed descriptions; the multi-level prompt word adjustment module 42 includes:

[0097] a judgment unit, configured to receive the initial prompt word, and judge whether the initial prompt word meets a preset level requirement, and obtain a judgment result;

[0098] a first result generating unit, configured to, if the judgment result is that the initial prompt word meets the preset hierarchical requirement, encode the initial prompt word into a recognizable condition vector of the generation model to generate the multi-level prompt word;

[0099] The second result generating unit is configured to, if the judgment result is that the initial prompt word does not meet the preset level requirement, grade and adjust the initial prompt word to generate the multi-level prompt word adapted to the generation model.

[0100] Furthermore, the second result generating unit includes:

[0101] a structure division unit, configured to, if the judgment result is that the initial prompt word does not meet the preset hierarchical requirement, parse the initial prompt word using a large language model, and perform structural division on the parsed initial prompt word according to the semantic hierarchy to obtain a divided prompt word;

[0102] a weight adjustment unit, configured to adjust the weights of the prompt words at each level in the divided prompt words to obtain adjusted prompt words;

[0103] A conditional vector generating unit is configured to encode the adjusted prompt word into a recognizable conditional vector of the generation model to generate the multi-level prompt word.

[0104] Furthermore, the target scene data generating module 43 includes:

[0105] a data splicing unit, configured to splice the multi-level prompt words and the pre-processed data to generate joint input data;

[0106] a data generation unit, configured to input the joint input data into the generative model, generate a latent space distribution of the generative model based on the multi-level prompt words through a conditional embedding layer, and generate scene attribute information of the target scene data based on the latent space distribution;

[0107] an alignment processing unit, configured to generate initial scene data based on the scene attribute information and the joint input data using a cross-modal attention mechanism, and perform multimodal information alignment processing on the initial scene data to generate aligned scene data;

[0108] A discrimination optimization unit is configured to generate the target scene data by performing discrimination optimization based on the aligned scene data through a discriminator of the generation model.

[0109] Furthermore, the data generating unit includes:

[0110] a feature fusion unit, configured to input the combined input data into the generative model, and perform feature fusion based on the identifiable conditional vectors of the multi-level prompt words and the feature map of the preprocessed data through a conditional embedding layer to adjust the spatial distribution;

[0111] A scene attribute information generating unit, configured to use the identifiable condition vector as an input channel of a generator in a generative adversarial network, so that the generator generates initial scene attribute information;

[0112] A denoising unit is used to gradually denoise the initial scene attribute information through a cross-attention layer of a diffusion model to generate the initial scene attribute information.

[0113] Furthermore, the target scene data generating module 43 further includes:

[0114] An evaluation result generating unit, configured to perform authenticity evaluation on the target scene data using a pre-trained discriminator network, and perform consistency evaluation on the target scene data using a cross-modal matching model, to obtain an evaluation result;

[0115] A parameter adjustment unit is used to adjust the model parameters of the generation model or the level weights of the multi-level prompt words according to the evaluation result.

[0116] To solve the above technical problems, the present application also provides a computer device. Figure 10 , Figure 10 This is a basic structural block diagram of the computer device in this embodiment.

[0117] The computer device 5 includes a memory 51, a processor 52, and a network interface 53 that are interconnected through a system bus. It should be noted that Figure 10Only a computer device 5 having three components, memory 51, processor 52, and network interface 53, is shown. However, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead. It should be understood by those skilled in the art that a computer device herein is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0118] Computer devices can be desktop computers, laptops, PDAs, cloud servers, etc. Computer devices can interact with users through keyboards, mice, remote controls, touchpads, or voice control devices.

[0119] The memory 51 includes at least one type of readable storage medium, including flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, etc. In some embodiments, the memory 51 may be an internal storage unit of the computer device 5, such as the hard disk or memory of the computer device 5. In other embodiments, the memory 51 may also be an external storage device of the computer device 5, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the computer device 5. Of course, the memory 51 may also include both the internal storage unit of the computer device 5 and its external storage devices. In this embodiment, the memory 51 is generally used to store the operating system and various application software installed on the computer device 5, such as the program code of the data generation method based on dynamic multi-level prompt words. In addition, the memory 51 can also be used to temporarily store various types of data that have been output or are to be output.

[0120] In some embodiments, the processor 52 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 52 is generally used to control the overall operation of the computer device 5. In this embodiment, the processor 52 is used to execute program code stored in the memory 51 or process data, such as executing the program code of the aforementioned method for generating data based on dynamic multi-level prompt words to implement various embodiments of the method for generating data based on dynamic multi-level prompt words.

[0121] The network interface 53 may include a wireless network interface or a wired network interface. The network interface 53 is generally used to establish a communication connection between the computer device 5 and other electronic devices.

[0122] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores a computer program. The computer program can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned data generation method based on dynamic multi-level prompt words.

[0123] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of each embodiment of the present application.

[0124] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.

Claims

1. A data generation method based on dynamic multi-level prompt words, characterized in that: include: receiving multimodal input data and preprocessing the multimodal input data to generate preprocessed data, wherein the multimodal input data includes environmental perception data and action sequence data; receiving an initial prompt word, and grading and adjusting the initial prompt word to generate a multi-level prompt word adapted to the generative model; The generation model performs data generation processing based on the pre-processed data and the multi-level prompt words to generate target scene data.

2. The data generation method based on dynamic multi-level prompt words according to claim 1 is characterized in that: The receiving multimodal input data and preprocessing the multimodal input data to generate preprocessed data includes: receiving the multimodal input data; voxelizing the point cloud data in the multimodal input data, performing resolution adjustment and normalization processing on the image data in the multimodal input data, and vectorizing the text data in the multimodal input data to generate initial processed data; The initial processed data is subjected to data enhancement to obtain enhanced data, and the enhanced data is subjected to data expansion by using a time series interpolation method to generate the pre-processed data.

3. The data generation method based on dynamic multi-level prompt words according to claim 1 is characterized in that: The multi-level prompt words have at least three levels, including overall categories, scene attributes, and detailed descriptions; the receiving of initial prompt words and grading and adjusting the initial prompt words to generate multi-level prompt words adapted to the generation model include: receiving the initial prompt word, and determining whether the initial prompt word meets the preset level requirement, and obtaining a determination result; If the judgment result is that the initial prompt word meets the preset level requirement, encoding the initial prompt word into a recognizable condition vector of the generation model to generate the multi-level prompt word; If the judgment result is that the initial prompt word does not meet the preset level requirement, the initial prompt word is graded and adjusted to generate the multi-level prompt word adapted to the generation model.

4. The data generation method based on dynamic multi-level prompt words according to claim 3 is characterized in that: If the judgment result is that the initial prompt word does not meet the preset level requirement, the initial prompt word is graded and adjusted to generate the multi-level prompt word adapted to the generation model, including: If the judgment result is that the initial prompt word does not meet the preset level requirement, the initial prompt word is parsed by the large language model, and the parsed initial prompt word is structurally divided according to the semantic level to obtain a divided prompt word; Adjusting the weights of the prompt words at each level in the divided prompt words to obtain adjusted prompt words; The adjusted prompt word is encoded into a recognizable condition vector of the generation model to generate the multi-level prompt word.

5. The data generation method based on dynamic multi-level prompt words according to claim 1 is characterized in that: The generating model performs data generation processing based on the pre-processed data and the multi-level prompt words to generate target scene data, including: Performing data splicing on the multi-level prompt words and the pre-processed data to generate joint input data; Inputting the combined input data into the generative model, generating a latent space distribution of the generative model based on the multi-level prompt words through a conditional embedding layer, and generating scene attribute information of the target scene data based on the latent space distribution; Using a cross-modal attention mechanism to generate initial scene data based on the scene attribute information and the joint input data, and performing multimodal information alignment processing on the initial scene data to generate aligned scene data; The target scene data is generated by performing discriminant optimization based on the aligned scene data through the discriminator of the generation model.

6. The data generation method based on dynamic multi-level prompt words according to claim 5, characterized in that: Inputting the combined input data into the generative model, generating a latent space distribution of the generative model based on the multi-level prompt words through a conditional embedding layer, and generating scene attribute information of the target scene data based on the latent space distribution, includes: Inputting the combined input data into the generative model, and performing feature fusion based on the identifiable conditional vectors of the multi-level prompt words and the feature map of the preprocessed data through a conditional embedding layer to adjust the spatial distribution; In a generative adversarial network, the identifiable condition vector is used as an input channel of a generator, so that the generator generates initial scene attribute information; The initial scene attribute information is gradually denoised through the cross attention layer of the diffusion model to generate the initial scene attribute information.

7. The data generation method based on dynamic multi-level prompt words according to any one of claims 1 to 6, characterized in that: After performing data generation processing based on the pre-processed data and the multi-level prompt words by the generation model to generate target scene data, the method further includes: Performing authenticity evaluation on the target scene data through a pre-trained discriminator network, and performing consistency evaluation on the target scene data through a cross-modal matching model to obtain an evaluation result; The model parameters of the generation model or the level weights of the multi-level prompt words are adjusted according to the evaluation result.

8. A data generation device based on dynamic multi-level prompt words, characterized in that: include: a data preprocessing module, configured to receive multimodal input data and preprocess the multimodal input data to generate preprocessed data, wherein the multimodal input data includes environmental perception data and action sequence data; a multi-level prompt word adjustment module, configured to receive initial prompt words, and classify and adjust the initial prompt words to generate multi-level prompt words adapted to the generation model; The target scene data generation module is used to perform data generation processing based on the pre-processed data and the multi-level prompt words through the generation model to generate target scene data.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the data generation method based on dynamic multi-level prompt words according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data generation method based on dynamic multi-level prompt words according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Large language model training method, system and equipment

    CN121256368A

  • Three-dimensional semantic scene completion method based on text semantic guidance

    CN121810954A