Deep learning-based data set enhancement method and system

By integrating cross-modal text encoders and syntactic analysis technology, stylized and semantically stable enhanced text is generated, which solves the problem of balancing semantic fidelity and expression diversity in existing technologies and achieves the controllability and diversity of text generation.

CN120705590AActive Publication Date: 2025-09-26JIAJIE TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511141956.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-09-26
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing text data augmentation technologies have difficulty balancing semantic fidelity and expression diversity, have weak style guidance capabilities, an uncontrollable generation process, and fail to effectively introduce non-verbal information for stylized expression, making it difficult to meet the requirements of complex tasks.

Method used

A pre-trained cross-modal text encoder is used to generate original semantic vectors, a projection function of multi-target modality is constructed, a global fused semantic vector is generated through syntactic analysis and mask fusion, a large language model is used for enhanced text generation, and non-linguistic information such as images and emotions are introduced for stylized control.

Benefits of technology

It achieves the generation of diverse and stylistically consistent enhanced text while maintaining semantic stability, improving the vividness and emotional appeal of the text, and is suitable for tasks such as creative writing and sentiment analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705590A_ABST
    Figure CN120705590A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data set enhancement, in particular to a data set enhancement method and system based on deep learning, and the method comprises the steps: inputting an original text into a pre-training cross-modal text encoder, and obtaining an original semantic vector; and presetting a multi-target mode, constructing a projection function of each mode, mapping the original semantic vector into a style guide vector of each target mode, and carrying out weighted fusion to generate a comprehensive style vector. Syntactic analysis is carried out on the original text, a word level mask is extracted, differential fusion is carried out on the context word vector and the comprehensive style vector based on the mask, and a global fusion semantic vector is generated. And mapping the text to an input space of a large language model through a trainable projection matrix, forming a soft prompt vector, injecting the soft prompt vector into a model input layer, and guiding to generate a plurality of enhanced texts with semantic loyalty and various styles to complete data set enhancement. According to the method and the device, the multi-stylization effect of the generated text is enhanced while the original semantic information is reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and more particularly to a dataset enhancement method and system based on deep learning. Background Art

[0002] With the widespread application of deep learning in natural language processing, high-quality, large-scale annotated data has become critical to improving model performance. However, in many practical application scenarios, such as healthcare, law, and finance, obtaining large amounts of annotated data is costly and time-consuming, leading to prominent problems such as sparse training data or class imbalance. Traditional data augmentation methods, such as synonym replacement, backtranslation, and random insertion / deletion, can increase data volume to a certain extent, but the generated text often lacks diversity and is prone to semantic drift, making it difficult to meet the dual requirements of semantic fidelity and expressive richness required for complex tasks.

[0003] In recent years, text augmentation techniques based on generative models have gradually emerged, especially the emergence of large language models (LLMs), which have provided powerful tools for generating high-quality text. However, directly using LLMs for augmentation often lacks control, resulting in highly random results that can easily deviate from the original semantics, creating "hallucinations" or irrelevant content.

[0004] In addition, most existing methods are limited to transformations within the text and fail to effectively introduce non-verbal information (such as vision, emotion, and hearing) to guide the stylized expression of the text, resulting in insufficient performance of enhanced text in terms of vividness, concreteness, and emotional appeal, making it difficult to meet the needs of tasks such as creative writing, advertising copywriting, and sentiment analysis that have specific requirements for expression style. Summary of the Invention

[0005] To address the technical issues in existing text data augmentation technologies, such as the difficulty in balancing semantic fidelity and expression diversity, weak style guidance capabilities, and uncontrollable generation processes, the present invention provides solutions in the following aspects.

[0006] In a first aspect, a dataset enhancement method based on deep learning includes: obtaining original text, inputting the original text into a pre-trained cross-modal text encoder to obtain an original semantic vector; presetting a set of multiple target modalities, constructing a projection function for each target modality, inputting the original semantic vector into each modality projection function, generating a style guidance vector for each target modality, and combining all target modality style guidance vectors to obtain a comprehensive style vector; performing syntactic analysis on the original text to extract a mask, and fusing the context word vector of the original text with the comprehensive style vector based on the mask to generate a global fused semantic vector; mapping the global fused semantic vector to the input embedding space of a pre-trained large language model via a trainable linear projection matrix to form a soft prompt vector, and injecting the soft prompt vector into the input layer of the large language model to generate multiple enhanced texts to complete the dataset enhancement.

[0007] Preferably, the pre-trained cross-modal text encoder is a pre-trained CLIP model, which consists of an image encoder and a text encoder.

[0008] Preferably, the set of multiple target modalities includes expression features of non-verbal information such as image target modality and emotional target modality.

[0009] Preferably, the comprehensive style vector includes: presetting a set of multiple target modalities to obtain a modal projection function corresponding to each target modality; according to the modal projection function corresponding to each target modality, mapping the original semantic vector to a function of the target modality semantic subspace, and outputting it as a style guidance vector; and performing weighted fusion on the style guidance vectors of different modalities to generate a comprehensive style vector.

[0010] Preferably, the modal projection function includes: taking the image modality as an example, the projection function can be modeled based on the directional characteristics of the cross-modal image encoder, and the image-text data pairs that are similar to the original text semantics in the image-text annotation data are obtained through manual screening according to the original text semantics. From the screened image-text data pairs, the cross-modal text encoder obtains the corresponding semantic vector and the cross-modal image encoder is used to obtain the corresponding image vector, and the image-text offset of each image-text data is obtained; the mean of the image-text offsets of all image-text data is obtained as the image style guidance direction, and the image style guidance direction is weightedly added to the original semantic vector of the original text to obtain the modal projection function.

[0011] Preferably, the global fusion semantic vector includes: performing syntactic analysis on the original text to extract the semantic stem vocabulary and non-semantic stem vocabulary of the original text, generating word-level masks based on the semantic stem vocabulary and non-semantic stem vocabulary of the original text; using the pre-trained language model encoding to obtain the context-aware word vector of each word in the original text, combining the context-aware word vector of each word with the comprehensive style vector of the original text, fusing according to the corresponding mask information, and obtaining the word-level fusion vector of each word after fusion; averaging and aggregating the word-level fusion vectors of each word to obtain the global fusion semantic vector.

[0012] Preferably, the generating of multiple enhanced texts includes the steps of: setting a projection matrix, multiplying the global fusion semantic vector by the projection matrix to obtain a soft prompt vector of the dimension required by the large language model input space, wherein the size of the projection matrix is ​​consistent with the dimension of the global fusion semantic vector and the size of the hidden layer dimension of the large language model; using the word segmenter of the large language model to segment the original text to obtain a word sequence, and obtaining the corresponding word embedding vector; splicing the soft prompt vector and the word embedding vector in the sequence dimension to form an extended input embedding sequence, wherein the soft prompt vector is at the front end of the extended input embedding sequence; the extended sequence is sent to the Transformer structure as the input of the large language model, and the large language model generates multiple enhanced texts in an autoregressive manner.

[0013] In a second aspect, a dataset enhancement system based on deep learning includes: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned dataset enhancement method based on deep learning is implemented.

[0014] The present invention has the following effects: 1. By introducing a pre-trained cross-modal text encoder to generate raw semantic vectors and constructing a learnable projection function for multiple target modalities, the raw semantic vectors are mapped into style guidance vectors for each modality. This achieves the effective transfer of cross-modal knowledge to text enhancement tasks. It can utilize the rich non-linguistic information contained in the cross-modal model to provide clear and quantifiable style guidance for text generation, thereby enhancing the expression effect of the generated text in different modal spaces.

[0015] 2. By performing dependency syntactic analysis on the original text, a mask vector is generated to distinguish semantic stem words from non-stem modifiers. Based on this mask, the context word vector and the comprehensive style vector are differentiated and fused. This achieves fine-grained semantic control of "stem freezing and edge perturbation", ensures the stability of the core semantics of the sentence, and avoids semantic drift caused by global style transfer. It also improves the expressive diversity and vividness of the generated text without destroying the original meaning. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a method flow chart of steps S1 to S4 in a dataset enhancement method based on deep learning in an embodiment of the present invention.

[0017] Figure 2 This is a structural block diagram of a dataset enhancement system based on deep learning in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.

[0019] Reference Figure 1 A dataset enhancement method based on deep learning includes steps S1 to S4, which are as follows: S1: Get the original text and input it into the pre-trained cross-modal text encoder to obtain the original semantic vector.

[0020] The input text samples to be augmented are read in batches from the original training set as the raw text. The raw text is preprocessed, including removing redundant spaces and standardizing punctuation to obtain standardized text.

[0021] The preprocessed original text is input into the pre-trained cross-modal text encoder to generate the corresponding original semantic vector. The original semantic vector represents the vector of the original text in the semantic space. The cross-modal text encoder adopts the text encoder in the CLIP (Contrastive Language-Image Pre-Training) model. The CLIP model is a well-known technical means and will not be described in detail.

[0022] S2: Preset a set of multiple target modalities, construct a projection function for each target modality, input the original semantic vector into each modality projection function, generate a style guidance vector for each target modality, and combine all target modality style guidance vectors to obtain a comprehensive style vector.

[0023] The original text Input the pre-trained cross-modal text encoder to obtain the corresponding original semantic vector .

[0024] After obtaining the original semantic vector corresponding to the original text, the original semantic vector represents the vector of the original text in the semantic space. Since different texts have different expression styles in multimodal scenarios, a preset target modality set is introduced to guide the original semantic vector to migrate towards a specific expression style, thereby achieving generalization of the content represented by the original text.

[0025] The target modality set includes but is not limited to image modality, emotional modality, etc., and each modality corresponds to an expression feature of non-verbal information.

[0026] For example, the image modality is used to enhance the spatial scene description ability and visual concreteness of the text; the emotional modality is used to strengthen the subjective emotional color and emotional intensity of the text.

[0027] By constructing a modal projection function that corresponds one-to-one to each target modality, a controllable stylized offset of the original semantic vector is achieved, thereby providing a guiding signal for the subsequent generation of diverse and stylistically consistent enhanced text.

[0028] Let the modal projection function It is defined as the original semantic vector Function mapped to the target modality semantic subspace, the output is the style guidance vector ,Right now .

[0029] The role of is not to completely change the original semantics, but to make moderate perturbations along the semantic direction defined by the target modality while maintaining the core semantic structure, so that the generated text can present the typical characteristics of the modality in terms of expression.

[0030] For example, when the target modality is "emotion", The output should be closer to semantic areas with high emotional intensity such as "excited", "sad" or "angry", thereby guiding the subsequent generation of more emotionally appealing text.

[0031] For different target modes, the construction method of the modal projection function can be designed differently according to the modal characteristics.

[0032] For example, taking the image modality as an example, the projection function Modeling can be performed based on the directional characteristics of the cross-modal image encoder. According to the semantics of the original text, the image-text data pairs with similar semantics to the original text in the image-text annotation data can be obtained by manual screening. From the screened image-text data pairs, the corresponding semantic vectors can be obtained using the cross-modal text encoder in step S1. , and use the cross-modal image encoder to obtain the corresponding image vector , get the image and text offset of each image and text data ; Get the average offset of all graphic data , as the image style guidance direction , and define , where the cross-modal image encoder uses the image encoder in the CLIP model. The CLIP model here is the same model as the CLIP model in step S1, so the cross-modal image encoder and the cross-modal text encoder in step S1 are trained using the same image-text dataset. is the adjustable gain coefficient, take the empirical value , to control the strength of style transfer, the number of image and text data that are similar in semantics to the original text is set to , which can be adjusted by implementers according to specific implementation scenarios.

[0033] For example, taking the emotional modality as an example, the projection function The trainable multi-layer perceptron (MLP) structure is used. Assume that the MLP consists of two fully connected layers and the activation function is , the overall structure is: ; During the training process, the original text is obtained through manual screening. Related text pairs with sentiment annotations ,in, yes A mood-enhanced version of It can be generated by artificial standard method, and the parameters are optimized by minimizing the loss function of mean square error. , , , is the optimization parameter that can be learned in MLP, and the original semantic text vector of the input is transformed into The corresponding output vector tends to the semantic area with high sentiment intensity.

[0034] When multimodal collaborative enhancement is enabled, multiple modal projection functions are used to generate style guidance vectors corresponding to the original text in different modalities, and a weighted fusion strategy is used to generate a comprehensive style vector: in, For the The weight value corresponding to a preset target mode, is the total number of preset target modes, For the original text A preset target modality style guidance vector, It indicates that the sum of the weights between different target modalities is 1, where the weight value can be adjusted by the implementer according to the implementation scenario. In the present invention, the weight values ​​between different target modalities are equal. By adjusting the weight values ​​between different target modalities, the weights of different target modalities in the comprehensive style vector can be adjusted.

[0035] The comprehensive style vector serves as a key input for subsequent semantic fusion. Its dimensions are consistent with the original semantic vector, ensuring compatibility with vector operations. The comprehensive style vector does not directly correspond to a specific text, but rather represents a "semantic trend" or "expressive tendency" and is used to introduce style perturbations at the level of non-core vocabulary.

[0036] S3: Perform syntactic analysis on the original text to extract a mask, and fuse the context word vector of the original text with the comprehensive style vector based on the mask to generate a fused semantic vector.

[0037] Through step S2, the comprehensive style vector corresponding to the original text can be obtained to achieve modal-level fusion. However, if the comprehensive style transfer is performed on the entire original text, it will cause semantic deviation problems in the original text. Therefore, word-level semantic analysis is performed on the original text to retain the main semantic information of the original text.

[0038] In order to achieve stable retention of the semantic trunk and stylized enhancement of non-stem components, the original text is subjected to dependency syntax analysis to identify its core grammatical structure, and a word-level mask vector is generated based on this. The mask value of the semantic trunk vocabulary in the word-level mask vector is set to 1, and the mask value of the non-semantic trunk vocabulary is set to 0, which is used to distinguish between semantic trunk vocabulary and non-stem vocabulary. Among them, dependency syntax is a well-known technical means and will not be repeated here.

[0039] On this basis, combined with the comprehensive style vector , perform differential fusion on the context-aware word vectors of each word, and finally generate a word-level fusion semantic vector that takes into account both semantic fidelity and expression diversity , providing high-quality semantic input for subsequent controllable text generation.

[0040] After obtaining the word-level mask vector, we obtain the context-aware word vector for each word in the original text. This vector is encoded using a pre-trained language model, such as BERT, RoBERTa, or sense2vec, and reflects the semantic information of the word in its specific context.

[0041] The original text The word is input into the pre-trained language model encoding to obtain the corresponding context-aware word vector ; Will With the style guidance vector generated above When performing fusion, a mask-based differentiated weighting strategy is adopted to implement the fusion principle of "main trunk freezing and edge perturbation". The specific fusion formula is as follows: For the words, and their fused word vectors Defined as: in, The main stem fusion coefficient is used to ensure that the main stem word vector mainly retains the original semantics and only introduces very slight style disturbance. The required value is small and the empirical value is taken. , which can be adjusted by implementers according to specific implementation scenarios; is the non-stem fusion coefficient, which allows non-stem words to fully absorb the information of the style guidance vector and realize the diversified expansion of expression methods. It requires a large value and takes the empirical value. , which can be adjusted by implementers according to specific implementation scenarios.

[0042] For the The mask value corresponding to each word in the word-level mask vector, Represented as the main semantic words in the original text, Represented as non-stem semantic words in the original text.

[0043] Original text The corresponding comprehensive style vector.

[0044] For the The context-aware word vector corresponding to each word.

[0045] This design ensures that the core semantic structure of the sentence is not destroyed, while the modifying components of non-main semantic information can be enhanced into more vivid or emotional expressions.

[0046] All word-level fusion vectors are aggregated into a global fusion semantic vector through an average operation: The vector The original text is retained The core semantic intent of the target modality is combined with the style guidance information of the target modality to form an enhanced semantic representation with rich semantics and controllable style. This vector will serve as a soft prompt input for the subsequent large language model generation process, guiding the model to generate enhanced text that is both faithful to the original meaning and has a diverse range of expression styles.

[0047] S4: Map the global fused semantic vector to the input embedding space of the pre-trained large language model via a trainable linear projection matrix to form a soft prompt vector, which is then injected into the input layer of the large language model to generate multiple enhanced texts and complete the dataset enhancement.

[0048] After obtaining the global fusion semantic vector generated in step S3 Finally, since the semantic space of the global fusion semantic vector and the large language model input embedding space is not consistent, a learnable soft prompt vector is formed to perform mapping alignment and then enhance text generation.

[0049] because The semantic space of the cross-modal encoder comes from the CLIP model, while the input embedding space of the large language model has a different dimensional structure. There is a mismatch between the spatial dimension and the semantic distribution between the two.

[0050] To this end, a trainable linear projection matrix is ​​introduced ,in, is the global fusion semantic vector dimension, is the hidden layer dimension of the large language model, for Through matrix multiplication, the global fusion semantic vector vf is mapped to the input space of the large language model: The resulting vector It is a soft hint vector, which does not correspond to any actual vocabulary, but carries the fused semantic and style information as the initial guiding signal for the generation process.

[0051] Use the word segmenter of the large language model to analyze the original text Perform word segmentation to obtain word sequence , The number of words in the word segmenter.

[0052] Get the corresponding word embedding vector .

[0053] Concatenate the soft hint vector h0 with the word embedding vector in the sequence dimension to form an extended input embedding sequence: The extended sequence is used as the input of the large language model and fed into the Transformer structure. Located at the front end of the sequence, the semantic information it carries is used as a context bias in the attention mechanism to participate in the calculation, thereby globally guiding the generation direction.

[0054] Large language models generate multiple augmented texts in an autoregressive manner ,in This is the preset number of generated items, with an empirical value of 1000. It can be adjusted by the implementer based on the specific implementation scenario.

[0055] The generation process can adopt a variety of decoding strategies, including: Top-k sampling: randomly sampling from the k words with the highest probability to increase diversity; Nucleus sampling (Top-p): sampling from the minimum set of words with a cumulative probability of p, balancing fluency and creativity; BeamSearch: retaining multiple candidate paths and selecting the output with the highest overall probability, suitable for scenarios that emphasize semantic coherence.

[0056] The system preferably sets a temperature coefficient (temperature≈0.7–0.9) to encourage diversity while maintaining fluency. The generated text is semantically consistent with the original text but reflects the style of the target modality.

[0057] For example, when the global fusion semantic vector is guided by the emotional modality, the generated text may emphasize subjective emotions more; when guided by the image modality, the text may contain more spatial descriptions and visual details.

[0058] Then it will As context guidance information, it is injected into the input layer of the large language model, driving the model to generate multiple enhanced texts with stable semantic information but diverse styles, completing the dataset enhancement.

[0059] The present invention also provides a data set enhancement system based on deep learning. Figure 2 As shown, the system includes a processor and a memory. The memory stores computer program instructions. When the computer program instructions are executed by the processor, a deep learning-based dataset enhancement method according to the first aspect of the present invention is implemented. The system also includes other components familiar to those skilled in the art, such as a communication bus and a communication interface. Their configuration and functions are well known in the art and are therefore not described in detail here.

[0060] It should be noted that those skilled in the art may make various modifications and improvements without departing from the scope of the present invention, and these modifications and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be based on the appended claims.

Claims

1. A dataset enhancement method based on deep learning, characterized in that: include: Get the original text and input it into the pre-trained cross-modal text encoder to obtain the original semantic vector; A set of multiple target modalities is preset, and a projection function is constructed for each target modality. The original semantic vector is input into each modality projection function to generate a style guidance vector for each target modality. The style guidance vectors of all target modalities are combined to obtain a comprehensive style vector. Performing syntactic analysis on the original text to extract a mask, and fusing the context word vector of the original text with the comprehensive style vector based on the mask to generate a global fused semantic vector; The global fusion semantic vector is mapped to the input embedding space of the pre-trained large language model via a trainable linear projection matrix to form a soft prompt vector. The soft prompt vector is injected into the input layer of the large language model to generate multiple enhanced texts and complete the dataset enhancement.

2. The method for data set enhancement based on deep learning according to claim 1, characterized in that: The pre-trained cross-modal text encoder is a pre-trained CLIP model, which consists of an image encoder and a text encoder.

3. The method for data set enhancement based on deep learning according to claim 1, characterized in that: The set of multiple target modalities includes expression features of non-language information such as image target modality and emotion target modality.

4. The method for data set enhancement based on deep learning according to claim 1, wherein: The comprehensive style vector includes: Preset a multi-target mode set and obtain the mode projection function corresponding to each target mode; According to the modal projection function corresponding to each target modality, the original semantic vector is mapped to the function of the target modality semantic subspace, and the output is the style guidance vector; The style guidance vectors of different modalities are weightedly fused to generate a comprehensive style vector.

5. The method for data set enhancement based on deep learning according to claim 4, characterized in that: The modal projection function includes: Taking the image modality as an example, the projection function can be modeled based on the directional characteristics of the cross-modal image encoder. According to the semantics of the original text, image-text data pairs with similar semantics to the original text are obtained in the image-text annotation data through manual screening. From the screened image-text data pairs, the cross-modal text encoder obtains the corresponding semantic vector and the cross-modal image encoder is used to obtain the corresponding image vector, and the image-text offset of each image-text data is obtained; the average image-text offset of all image-text data is obtained as the image style guidance direction, and the image style guidance direction is weightedly added to the original semantic vector of the original text to obtain the modal projection function.

6. The method for data set enhancement based on deep learning according to claim 1, characterized in that: The global fusion semantic vector includes: Performing syntactic analysis on the original text to extract the semantic stem words and non-semantic stem words of the original text, and generating word-level masks based on the semantic stem words and non-semantic stem words of the original text; The pre-trained language model is used to encode the context-aware word vector of each word in the original text. The context-aware word vector of each word is combined with the comprehensive style vector of the original text and fused according to the corresponding mask information to obtain the word-level fusion vector of each word after fusion. The word-level fusion vectors of each word are averaged and aggregated to obtain the global fusion semantic vector.

7. The method for data set enhancement based on deep learning according to claim 1, characterized in that: The method of generating multiple enhanced texts comprises the steps of: Set the projection matrix and multiply the global fusion semantic vector by the projection matrix to obtain a soft prompt vector of the dimension required by the large language model input space. The size of the projection matrix is ​​consistent with the dimension of the global fusion semantic vector and the hidden layer dimension of the large language model. Use the word segmenter of the large language model to segment the original text, obtain word sequence, and obtain the corresponding word embedding vector; Concatenate the soft hint vector and the word embedding vector in the sequence dimension to form an extended input embedding sequence, where the soft hint vector is at the front of the extended input embedding sequence; The extended sequence is fed into the Transformer structure as the input of the large language model, and the large language model generates multiple enhanced texts in an autoregressive manner.

8. A dataset enhancement system based on deep learning, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a data set enhancement method based on deep learning according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Hierarchical cross-modal sentiment analysis method based on text guidance and related device

    CN118821054A

  • Unsupervised cross-modal hash retrieval method based on CLIP and attention fusion mechanism

    CN118861327A

  • Unmanned aerial vehicle vector formation cooperative control method and system

    CN119882828A

  • Semantic comprehension driven cross-modal information fusion and retrieval method and system

    CN120448563A

  • Multimodal financial technology deep learning core with joint optimization of vector-quantized variational autoencoder and neural upsampler

    US12327190B1