Data enhancement method for event information deduction

Through the data enhancement method for event information deduction, and using technical means such as generating adversarial networks and event graphs, the problem of scarcity of training data is solved, the learning effect and generalization ability of event information deduction model are improved, and more accurate and efficient information deduction and decision support are achieved.

CN120180378AActive Publication Date: 2025-06-20BEIJING INST OF COMP TECH & APPL
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510675888.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

In intelligence processing, due to the special nature of certain fields, including sensitivity and confidentiality, available training data is scarce and difficult to obtain, limiting the application effect of machine learning algorithms, especially deep learning models, in intelligence analysis.

Method used

A data augmentation method for event information deduction is proposed. By generating an adversarial network, the original data set is expanded and more diverse training samples are generated, including embedding network E, recovery network R, generation network G, discriminant network D and supervision network S. The event graph and multimodal embedding representation are used for data fusion and characterization description, the LSTM structure is introduced to capture time dependencies, and a multi-scale feature extraction mechanism is adopted and a new strategy combining adversarial loss and supervision guidance is adopted.

Benefits of technology

It effectively improves the ability to deduce information in small samples, improves the learning effect and generalization ability of the event deduction model, ensures the consistency and logical consistency of the generated data points on the timeline, and enhances the model's ability to deduce future events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180378A_ABST
    Figure CN120180378A_ABST
Patent Text Reader

Abstract

The invention relates to a data enhancement method for event information deduction, and belongs to the field of artificial intelligence. In order to solve the problem that available training data are scarce and difficult to obtain, data fusion is performed on collected multi-source heterogeneous data by using an event graph, unified representation description is performed on the multi-source heterogeneous data by using multi-modal embedding representation of a fusion network to obtain fused time sequence data, and then the time sequence data are subjected to data fusion. And training by using the generative adversarial network, so that the generative network G generates an implicit vector Z based on random noise, and the recovery network R converts the implicit vector Z back to new time sequence data. According to the method, the generative adversarial network is adopted to expand the original data set to generate more diversified training samples, so that the learning effect and generalization ability of the event deduction model are improved, and the small sample event information deduction ability is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and specifically relates to a data augmentation method for event information deduction. Background Art

[0002] In intelligence processing, the demand for dynamic information analysis and processing of document data is constantly increasing. However, due to the special nature of certain fields, including sensitivity and confidentiality, the available training data is scarce and difficult to obtain, which greatly limits the application effect of machine learning algorithms, especially deep learning models, in intelligence analysis. To address this challenge, the present invention proposes a data augmentation method for event information deduction, conducts research on generative adversarial networks for time-series data to achieve data augmentation, thereby effectively solving the small-sample problem and supporting the construction of a more accurate and efficient information deduction and decision support system. Summary of the Invention

[0003] (I) Technical Problems to be Solved The technical problem to be solved by the present invention is how to provide a data augmentation method for event information deduction to solve the problem of scarce and difficult-to-obtain available training data.

[0004] (II) Technical Solutions To solve the above technical problems, the present invention proposes a data augmentation method for event information deduction, which includes the following steps: Step 1: Data Collection Collect multi-dimensional intelligence data from different data sources, perform preliminary processing on the intelligence data to obtain multi-source heterogeneous data, and load it into a unified multi-modal database; Step 2: Data Processing Use an event graph to perform data fusion on the collected multi-source heterogeneous data, use the multi-modal embedding representation of the fusion network to perform a unified characterization description of the multi-source heterogeneous data, obtain the fused time-series data, and then perform data standardization and normalization; Step 3: Model Training The generative adversarial network includes: an embedding network E, a recovery network R, a generative network G, a discriminant network D, and a supervision network S. By training the generative adversarial network, the generative network G generates a latent vector Z based on random noise, and the recovery network R converts the latent vector Z back into new time-series data; Step 4: Data Augmentation After the model training is completed, provide random noise to the generative network G to generate a latent vector Z, and use the recovery network R to convert the latent vector Z back into new time-series data.

[0005] (III) Beneficial Effects The present invention proposes a data augmentation method for event information deduction. The present invention discloses a data augmentation method for event information deduction, and its main advantages are reflected in the following aspects: (1) The present invention uses an event graph to fuse multi-source heterogeneous data collected, and uses the multi-modal embedding representation of the fusion network to uniformly represent and describe the multi-source heterogeneous data, thereby realizing the unification of multi-source heterogeneous data; (2) In order to accurately capture the time-dependent relationship of events in intelligence, an LSTM structure is introduced into the traditional generation network to learn and reproduce the complex dynamic patterns in time series data, ensuring the coherence and logical consistency of the generated data points on the time axis.

[0006] (3) In order to extract the information of multi-level intelligence, a multi-scale feature extraction mechanism is adopted in the discriminant network to adapt to data requirements of different granularities, enabling the discriminant network to not only evaluate the authenticity of the overall sequence, but also perform fine-grained discrimination on local features, improving the discrimination accuracy.

[0007] (4) A new strategy combining adversarial loss and supervision guidance is proposed, enabling the generation network to predict possible future development paths while imitating history, thereby enhancing the model's ability to deduce future events.

[0008] (5) The present invention uses a generative adversarial network to expand the original data set to generate more diverse training samples, thereby improving the learning effect and generalization ability of the event deduction model, and effectively enhancing the small-sample event information deduction ability. Description of the Drawings

[0009] Figure 1 is the flowchart of the method of the present invention. Detailed Embodiments

[0010] To make the objectives, content and advantages of the present invention clearer, the following further describes in detail the specific embodiments of the present invention with reference to the drawings and embodiments.

[0011] The present invention aims to solve the problem of insufficient accuracy of event deduction in the prior art due to the lack of sufficient training samples. Specifically, for the small-sample problem in the event deduction scenario, the present invention uses a generative adversarial network to expand the original data set to generate more diverse training samples, thereby improving the learning effect and generalization ability of the event deduction model.

[0012] The present invention proposes a data augmentation method for event information deduction, which is used to solve the small-sample event deduction task and effectively enhance the small-sample event information deduction ability.

[0013] The present invention provides a data augmentation method for event information deduction, and the overall process is asFigure 1 As shown, generally speaking, it includes four steps, namely data collection, data processing, model training, and data augmentation.

[0014] Collect multi-dimensional intelligence data from different data sources, conduct preliminary processing on the intelligence data to obtain multi-source heterogeneous data, and load it into a unified multi-modal database.

[0015] Specifically, first, by integrating data sources (i.e., data sources such as official document systems and open-source data), and introducing high-concurrency real-time data streams, capture multi-dimensional intelligence, set up a temporary storage buffer for preliminary storage and preprocessing of a large amount of received raw data, conduct preliminary filtering and format standardization on the data in the buffer, use ETL tools to extract and transform data from different sources to obtain multi-source heterogeneous data, and load it into a unified multi-modal database.

[0016] Use an event graph to perform data fusion on the collected multi-source heterogeneous data, use the multi-modal embedding representation of the fusion network to perform unified characterization and description of the multi-source heterogeneous data, obtain the fused time-series data, and then perform data standardization and normalization.

[0017] S21. First, construct an event graph. The event nodes in the event graph are composed of multi-source heterogeneous data (images, texts...), that is, the event graph is a heterogeneous network constructed for multi-modal data. Use the DeepWalk algorithm to perform embedding learning on the event nodes in the event graph (each node represents an event), and generate an event node sequence (such as [A, B, C]) through a truncated random walk algorithm. Use the co-occurrence relationship of the event nodes in the sequence to represent its local neighbor structure (such as node B is a neighbor of A, and C is a neighbor of B), and then apply the Word2Vec word embedding model to the obtained sequence to obtain the network representation of the event nodes.

[0018] For example, assume there is an action-related event, and the event nodes include: 1. Intelligence collection (data types: text, image): For example, obtain deployment information through satellite images and intelligence reports.

[0019] 2. Action decision (data type: timestamp): For example, the time point for deciding whether to initiate an action.

[0020] 3. Resource investment (data type: structured data): For example, the number of personnel and equipment invested.

[0021] Random walk sequence: [Intelligence collection, Resource investment, Action decision], and the nodes randomly walk to form a sequence. An event graph is composed of event nodes, and the event graph is defined based on all event nodes. The multi-modal data is equivalent to the values in the event nodes.

[0022] S22. For the heterogeneous network constructed for multimodal data, first construct an embedding representation for each modality; text data can output a vector representation through the TF-IDF method, and image data can obtain a vector representation through a convolutional network and then through a fully connected layer. Then, map the embedding representations of different modalities to an embedding space of the same dimension. In each iteration, maximize the similarity between the embedding representations of linked nodes and minimize the similarity of unlinked nodes. By continuously adjusting the parameters in the vector representation generation process, finally obtain the multimodal embedding representation of the fusion network.

[0023] The fusion network comes after the heterogeneous network. It means that after the multimodal embedding representation (text data outputs a vector representation through the TF-IDF method, and image data obtains a vector representation through a convolutional network and then through a fully connected layer) generated by the heterogeneous network, through the fusion network, map the embedding representations of different modalities to an embedding space of the same dimension, and finally obtain the multimodal embedding representation of the fusion network.

[0024] S23. Arrange the multimodal embedding representations of the fusion network in chronological order to obtain the fused time series data X = {x1, x2,..., x T}, where each x t is the multimodal embedding representation of the fusion network at time t. Perform standardization and normalization processing on these data so that all features have the same numerical range [-1, 1], which is beneficial for the model to learn the relative relationship between data and generate standardized time series data.

[0025] The generative adversarial network includes: an embedding network E, a recovery network R, a generative network G, a discriminative network D, and a supervision network S. By training the generative adversarial network, the generative network G generates a latent vector Z based on random noise, and the recovery network R converts the latent vector Z back to new time series data; S31. First, perform the pre-training stage: S311. Pre-train the embedding network E and the recovery network R. Use the processed standardized time series data X as the dataset to train the embedding network E and the recovery network R. The purpose is to enable the recovery network R to accurately reconstruct the original time series from the latent vector output by the embedding network E. The goal of this stage is to minimize the reconstruction loss .

[0026] S312. Pre-train the generative network G and the supervision network S. The random noise z tries to generate a latent vector , while the supervision network S tries to predict whether these latent vectors come from the embedding network E. At this time, the discriminative network D is not involved. The purpose is to let the generation network G learn to generate latent vectors close to the output of the embedding network E. The loss function for pre-training is .

[0027] Introduce the LSTM structure in the generation network G to learn and reproduce the complex dynamic patterns in the time series data, ensuring the coherence and logical consistency of the generated data points on the time axis.

[0028] S32. Secondly, perform the adversarial training phase: S321. Update the generation network G and the supervision network S. Fix the other parts and only update the generation network G and the supervision network S so that G can generate more realistic latent vectors and S can better predict the authenticity of these latent vectors. In this process, G is not only affected by the feedback from the discriminative network D, that is, the adversarial loss, but also by the additional guidance from S, that is, the supervision loss; the input of the discriminative network D is the latent vector from the embedding network E or the generation network G (the real sample from E or the sample generated by G), and the output is the probability of sample authenticity. The total loss function is:

[0029] where γ is a hyperparameter used to adjust the balance between the adversarial loss and the supervision loss . The adversarial loss and the supervision loss are as follows:

[0030] where, is the random noise distribution, D is the discriminative network, denotes the expectation, and G(z) represents the time series data generated by the generation network G based on the input noise z. is the latent vector after the real time series data is transformed by the embedding network, and S is the supervision network.

[0031] S322. Update the discriminative network D. Fix the generation network G and other components and only update the discriminative network D to make it more effective in distinguishing true and false time series. This step is achieved by minimizing the discriminative loss.

[0032] Adopt a multi-scale feature extraction mechanism in the discriminative network D to adapt to different granularity data requirements, enabling the discriminative network D to not only evaluate the authenticity of the overall sequence but also perform fine-grained discrimination on local features, improving the discrimination accuracy.

[0033] The discrimination loss of the discrimination network D is usually to maximize the discrimination accuracy of real samples and generated samples (the smaller the loss, the better the discrimination network is trained, so the accuracy is high), while the adversarial loss of the generation network G is to deceive the discrimination network into thinking that the generated samples are real. Therefore, in step S322, the discrimination loss refers to the loss function of the discrimination network, and the adversarial loss is a part of the loss of the generation network. The two together constitute the process of adversarial training.

[0034] S323. Update the embedding network E and the recovery network R. Adjust the embedding network E and the recovery network R again. By simultaneously using the real latent vector E(X) and the generated latent vector G(z), force the recovery network R to maintain a high-precision reconstruction ability when facing real or generated latent vectors. The loss function at this stage is:

[0035] where β is a hyperparameter used to adjust the balance between the reconstruction loss and the supervision loss. The reconstruction loss and the supervision loss are specifically as follows:

[0036] where, denote E(X) and G(z). E(X) is the latent vector after the real time series data is transformed by the embedding network, and G(z) represents the latent vector generated by the generation network based on the input noise z.

[0037] S33. Then, enter the alternating optimization stage: In each iteration, update the parameters of each network in the order of S321 - S323 above to ensure that each network can be fully trained without overfitting. This alternating update strategy helps to stabilize the entire training process. Adjust hyperparameters such as the learning rate and batch size of the model according to experiments to obtain better generation effects. When the model reaches the predetermined stop condition (such as the maximum number of iterations T max , the discrimination network cannot distinguish the authenticity of the generated data), end the training process.

[0038] After the model training is completed, provide random noise to the generation network G to generate the latent vector Z, and use the recovery network R to convert the latent vector Z back to new time series data. These newly generated time series data should be as close as possible to the actual observed data distribution, but not a simple copy and paste. The specific operations are as follows: S41. Prepare the random noise z: Extract a set of random numbers from a predefined noise distribution (such as the standard normal distribution or the uniform distribution) as the input.

[0039] S42. Generate the latent vector: Input the random noise z into the generation network G to obtain the latent vector Z = G(z).

[0040] S43. Convert to time series: Use the recovery network R to convert the latent vector Z back into the time series form X ^ = R(Z). At this time, X ^ is the newly generated time series data. To ensure the quality of the generated data, multiple sets of samples are usually generated, and those closest to the true distribution are selected from them. In addition, evaluation metrics (such as KL divergence, Wasserstein distance, etc.) can be used to quantify the similarity between the generated data and the true data.

[0041] S44. Post - processing: The generated time series data often undergoes standardization or normalization processing, so inverse standardization / inverse normalization is required to restore its original numerical range.

[0042] The augmented data generated after post - processing can be used as the input of the subsequent event information deduction model, thereby improving the training quality of the event information deduction model.

[0043] The present invention discloses a data augmentation method for event information deduction, and the main advantages are reflected in the following aspects: (1) The present invention uses an event graph to fuse multi - source heterogeneous data collected, and uses the multi - modal embedding representation of the fusion network to uniformly represent and describe the multi - source heterogeneous data, thereby achieving the unification of multi - source heterogeneous data; (2) In order to accurately capture the time - dependent relationship of events in intelligence, an LSTM structure is introduced into the traditional generation network to learn and reproduce the complex dynamic patterns in the time series data, ensuring the coherence and logical consistency of the generated data points on the time axis.

[0044] (3) In order to extract the information of multi - level intelligence, a multi - scale feature extraction mechanism is adopted in the discriminant network to adapt to data requirements of different granularities, so that the discriminant network can not only evaluate the authenticity of the overall sequence, but also perform fine - grained discrimination on local features, improving the discrimination accuracy.

[0045] (4) A new strategy combining adversarial loss and supervised guidance is proposed, enabling the generation network to predict possible future development paths while imitating the history, thereby enhancing the model's ability to deduce future events.

[0046] (5) The present invention uses a generative adversarial network to expand the original data set to generate more diverse training samples, thereby improving the learning effect and generalization ability of the event deduction model, and effectively enhancing the ability of small - sample event information deduction.

[0047] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A data augmentation method for event information deduction, characterized in that, The method includes the following steps: Step 1: Data collection Collect multi-dimensional intelligence data from different data sources, preliminarily process the intelligence data to obtain multi-source heterogeneous data, and load it into a unified multi-modal database; Step 2: Data processing Use an event graph to perform data fusion on the collected multi-source heterogeneous data, and use the multi-modal embedding representation of the fusion network to perform unified characterization and description of the multi-source heterogeneous data to obtain the fused time series data, and then perform data standardization and normalization; Step 3: Model training The generative adversarial network includes: an embedding network E, a recovery network R, a generative network G, a discriminative network D, and a supervision network S. By training the generative adversarial network, the generative network G generates a latent vector Z based on random noise, and the recovery network R converts the latent vector Z back into new time series data; Step 4: Data augmentation After the model training is completed, provide random noise to the generative network G to generate a latent vector Z, and use the recovery network R to convert the latent vector Z back into new time series data.

2. The data augmentation method for event information deduction according to claim 1, characterized in that, The said Step 1 includes: by integrating data sources and introducing high-concurrency real-time data streams, capturing multi-dimensional intelligence, setting up a temporary storage buffer for initially storing and preprocessing a large amount of received raw data, preliminarily filtering and format standardizing the data in the buffer, using an ETL tool to extract and transform data from different sources to obtain multi-source heterogeneous data, and loading it into a unified multi-modal database.

3. The data augmentation method for event information deduction according to claim 1, characterized in that, The said Step 2 includes: S21. First, construct an event graph, where the event nodes in the event graph are composed of multi-source heterogeneous data; use the DeepWalk algorithm to perform embedding learning on the event nodes in the event graph, and generate an event node sequence through a truncated random walk algorithm, use the co-occurrence relationship of the event nodes in the sequence to represent its local neighbor structure, and then apply the Word2Vec word embedding model to the obtained sequence to obtain the network representation of the event nodes; S22. For the heterogeneous network constructed for multi-modal data, first construct an embedding representation for each modality, and then map the embedding representations of different modalities to an embedding space of the same dimension. In each iteration, maximize the similarity between the embedding representations of linked nodes and minimize the similarity between unlinked nodes. By continuously adjusting the parameters in the vector representation generation process, finally obtain the multi-modal embedding representation of the fusion network; S23. The multi-modal embedding representations of the fusion network are arranged in chronological order to obtain the fused time series data. , where each x t is the multi-modal embedding representation of the fusion network at time t; these data are standardized and normalized so that all features have the same numerical range [-1, 1], generating standardized time series data.

4. The data augmentation method for event information deduction according to claim 3, characterized in that, The text data outputs a vector representation through the TF-IDF method.

5. The data augmentation method for event information deduction according to claim 3, characterized in that, The image data is processed by a convolutional network and then obtains a vector representation through a fully connected layer.

6. The data augmentation method for event information deduction according to claim 3, characterized in that, The said Step 3 includes: S31. First, perform the pre-training stage: S311, pre-trained embedding network E and recovery network R; using the processed and standardized time series data X as a dataset to train the embedding network E and the recovery network R, enabling the recovery network R to accurately reconstruct the original time series from the hidden vector output by the embedding network E ; The goal of this stage is to minimize the reconstruction loss ; S312, pre-training generation network G and supervision network S; random noise z attempts to generate a latent vector through generation network G , while supervision network S endeavors to predict whether these latent vectors come from embedding network E; at this time, discriminative network D is not involved, and the aim is to enable generation network G to learn to generate latent vectors close to the output of embedding network E; the loss function for pre-training is ; S32. Second, perform the adversarial training stage: S321. Update the generation network G and the supervision network S; fix the other parts and only update the generation network G and the supervision network S so that G can generate more realistic latent vectors and S can better predict the authenticity of these latent vectors; during this process, G is not only affected by the feedback from the discriminative network D, i.e., the adversarial loss, but also by the additional guidance from S, i.e., the supervision loss; the input of the discriminative network D is the latent vector from the embedding network E or the generation network G, and the output is the probability of sample authenticity. The total loss function is: where γ is a hyperparameter used to adjust the balance between the adversarial loss and the supervised loss ; S322. Update the discriminative network D; fix the generation network G and other components and only update the discriminative network D to enable it to more effectively distinguish between true and false time series; this step is achieved by minimizing the discriminative loss; S323. Update the embedding network E and the recovery network R; readjust the embedding network E and the recovery network R again. By simultaneously using the true latent vector E(X) and the generated latent vector G(z), force the recovery network R to maintain a high-precision reconstruction ability when facing either the true or the generated latent vector; the loss function at this stage is: where β is a hyperparameter used to adjust the balance between the reconstruction loss and the supervision loss. The reconstruction loss and the supervision loss are specifically as follows: Among them, denote E(X) and G(z), where E(X) is the latent vector after the real time series data is transformed by the embedding network, and G(z) represents the latent vector generated by the generation network based on the input noise z; S33. Then, enter the alternating optimization stage: In each iteration, update the parameters of each network in the order of S321 - S323 above to ensure that each network can be fully trained without overfitting; when the model reaches the predetermined stopping condition, end the training process.

7. The data enhancement method for event information deduction according to claim 6, wherein, In S312, an LSTM structure is introduced into the generation network G to learn and reproduce the complex dynamic patterns in the time series data, ensuring the coherence and logical consistency of the generated data points on the time axis.

8. The data enhancement method for event information deduction according to claim 6, wherein, In S321, the adversarial loss and the supervision loss are specifically as follows: Among them, is the random noise distribution, D is the discriminative network, denotes the expectation, while G(z) represents the time series data generated by the generative network G based on the input noise z; is the latent vector after the real time series data is transformed by the embedding network, and S is the supervision network.

9. The data enhancement method for event information deduction according to claim 6, wherein, In S322, a multi-scale feature extraction mechanism is adopted in the discriminative network D to adapt to data requirements at different granularities, enabling the discriminative network D to not only evaluate the authenticity of the overall sequence but also perform fine-grained discrimination on local features.

10. The data enhancement method for event information deduction according to claim 6, wherein, The specific steps of Step Four include: S41. Prepare the random noise z: Draw a set of random numbers from a predefined noise distribution as the input; S42. Generate the latent vector: Input the random noise z into the generation network G to obtain the latent vector Z = G(z); S43. Convert to time series: Use the recovery network R to convert the latent vector Z back into the time series form X ^ = R(Z); At this time, X ^ is the newly generated time series data; S44. Post-processing: Perform denormalization and inverse normalization on the generated time series data to restore its original numerical range.

Citation Information

Patent Citations

  • Serialized micro-service resource prediction method based on adversarial learning and heterogeneous graph learning

    CN117648197A

  • Micro-application system anomaly detection method and device based on graph neural network

    CN117972542A

  • Semantic search method and related device

    CN117992585A

  • Fault classification method based on Nadam-TimeGAN and MCNN-GRU

    CN118568553A

  • Intelligent data enhancement method and device based on generative adversarial network and multi-modal data, and medium

    CN118606715A