Key attribute extraction method and system based on multi-modal normalization
Patent Information
- Application Number
- CN202510688784.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-26
Smart Images

Figure CN120705539A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and data analysis technology, and specifically to a key attribute extraction method and system based on multimodal normalization. Background Art
[0002] With the deepening of digital transformation, data across various industries is presenting in multiple, heterogeneous forms, including structured (database tables), semi-structured (XML, JSON), and unstructured (text, images, audio, video, etc.). In areas such as data security, network incident tracing, and intelligence analysis, comprehensive analysis requires accurate extraction of key attribute features from various data types to achieve data alignment, integration, and effective utilization.
[0003] With the rapid advancement of information technology and digital transformation, the scale of data collection and processing across various industries has rapidly expanded. Data sources are becoming increasingly diverse, and data forms and structures are becoming increasingly complex and diverse, gradually exhibiting distinct multi-source heterogeneity. Multi-source heterogeneous data primarily refers to data sets from different sources and channels, with varying formats and structural characteristics. In application areas such as cybersecurity, data security incident tracing and analysis, intelligence analysis, and threat warning, there is an urgent need to integrate and conduct in-depth correlation analysis of this multimodal, multi-source, and heterogeneous data to support risk prediction, incident tracking, and decision-making. However, a key prerequisite for achieving this goal is the accurate and efficient extraction of representative key attribute features from these diverse data sources to enable subsequent data fusion and unified analysis.
[0004] Currently, the mainstream technology for extracting key attributes from multi-source, heterogeneous data relies primarily on traditional natural language processing (NLP) methods and specialized deep learning-based models. Specifically, traditional NLP methods, such as rule matching, dictionary search, and conditional random field (CRF) models, rely heavily on manually formulated rules and large amounts of annotated data, resulting in poor generalization and significant model maintenance costs. These methods struggle to adapt quickly to new data patterns and terminology, leading to extremely low processing efficiency. Deep learning-based methods, such as CNNs, RNNs, LSTMs, and pre-trained models (Transformers), while performing well for specific tasks, are generally limited by the requirement for large amounts of manually annotated data for model training, making them ineffective in generalizing to new data sources and attribute categories. Furthermore, the accuracy of these models declines significantly when data features complex modalities, variable structures, and the frequent presence of non-standard terminology and symbols. Furthermore, the models often need to be completely rebuilt with each new data type or scenario, severely limiting their flexibility and scalability.
[0005] Based on the above background and needs, the present invention innovatively designs a key attribute extraction method and system based on multimodal normalization, which thoroughly solves the limitations and shortcomings of the existing technology with efficient, accurate and highly generalized automation technology, and effectively promotes the further development and practical application of data analysis technology. Summary of the Invention
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a key attribute extraction method based on multimodal normalization, specifically comprising the following steps: Step S1: Implementation environment and data preparation; Build an implementation environment within the network security enterprise intranet; Server environment: Two Tesla A100 GPU servers running Python, PyTorch, and the Transformers deep learning framework. Data storage: MongoDB, Elasticsearch, and MinIO are used to build data warehouses and feature databases; Data sources: Threat intelligence platform real-time API interface, network intrusion detection system (IDS) data interface, and terminal log real-time collection interface.
[0007] Step S2: data collection and preliminary processing; Collect unstructured text data from the threat intelligence platform API in real time, and collect network traffic records and terminal system logs from IDS devices; Remove missing records and redundant data, automatically mark data types, and uniformly store data records in JSON format to form a standardized data foundation.
[0008] Step S3, normalization processing of polymorphic elements; Based on the initial data records, a polymorphic element normalizer is deployed to initially establish an expert knowledge base and an automatically updated terminology standardization library to convert text data into unified data.
[0009] Step S4: Simplify the semantics of complex text and construct a UmVAE model; Step S5: feature extraction; Based on the simplified sentences, a feature extractor is deployed to embed the Transformer model features into the semantic vector of the generated data; the DBSCAN unsupervised clustering method is used to automatically identify candidate features, and automatic classification is carried out in combination with the expert knowledge rule base.
[0010] Preferably, step S4 constructs the UmVAE model, specifically comprising the following steps: Step A1: input normalized data after normalization of polymorphic elements; Step A2: Using unsupervised training, reconstruction error and KL divergence optimization model parameter reparameterization technique to obtain the latent space vector; Step A3: Implement model inference: Input the above normalized data and simplify it through UmVAE to obtain simplified data with precise semantics and clear structure.
[0011] A key attribute extraction system based on multimodal normalization includes a data acquisition unit, a data processing unit and a feature extraction unit; Among them, the data acquisition unit is composed of a real-time data acquisition unit and a data preprocessing unit; The data collection unit collects factual data from the security intelligence platform and threat intelligence database, network security monitoring equipment, and terminal equipment and server system logs, and cleans and removes noise from the data; The data preprocessing unit includes preliminary processing of data format normalization, preliminary marking of data labels, and data indexing and storage; Data format normalization is a preliminary process that stores heterogeneous data in a unified format to facilitate subsequent unified processing. It extracts the core information of the original data and determines the unified field name and data structure. Preliminary data labeling Preliminary data labels are automatically generated during the preprocessing stage; Data indexing and storage uses an efficient data indexing engine to index processed data for fast query. Data storage strategy: It is recommended to store structured data in a time series database and unstructured data in an object storage. Clear storage paths and naming conventions are defined to ensure data orderliness and efficient retrieval capabilities. The data processing unit includes a pattern matching and replacement module, an automatic construction and iteration knowledge mapping base module, an intelligent abbreviation and non-standard terminology recognition module, and a complex text simplification processing unit; Among them, the pattern matching and replacement module builds a regular expression library to identify non-standard characters; The automatic construction and iteration of the knowledge mapping library module identifies newly emerging non-standard terms and abbreviations from continuously collected data, automatically matches and tags them, and automatically discovers new terms and abbreviations using unsupervised methods. When the frequency of new terms exceeds a threshold, it automatically updates the knowledge base and automatically prompts domain experts for verification and update. After confirmation by domain experts, the knowledge base is automatically updated. The non-standard terminology recognition module uses the BERT pre-trained language model combined with conditional random fields (CRF) to automatically identify abbreviations and non-standard terms. The pre-trained language model is used to encode contextual features of text input; sequence labeling is performed through the CRF model to identify non-standard terms and abbreviations in the text.
[0012] The complex text simplification processing unit is used to build the UmVAE model and train and optimize the constructed UmVAE model; Among them, the UmVAE model consists of three sub-modules: multimodal variational encoder, semantic latent space, and variational decoder; The multimodal variational encoder inputs the binary normalized standard text sequence and uses the Transformer pre-trained language model to convert the input text into a continuous dense vector. The semantic latent space is essentially a low-dimensional dense semantic vector space that can continuously and smoothly represent the semantics of text. During model training, the KL divergence regularization term is used to ensure the structural stability and generalization of the latent semantic space, enabling efficient mapping and conversion from complex text to clear semantics. The variational decoder uses Transformer as the decoder network architecture, takes the vector in the latent space as input, and maps it back to the data text with clear semantics and standardized structure; The feature extraction unit includes feature embedding and candidate feature identification, expert knowledge rule base construction, candidate feature classification and filtering mechanism; Among them, feature embedding and candidate feature identification are specifically as follows: The input text comes from the standardized text output of the complex text simplifier. A pre-trained Transformer model is used to convert text sentences into contextual semantic feature embedding vectors. An unsupervised semantic clustering algorithm is then used to perform cluster analysis on the embedded features. Frequently occurring phrases, short sentences, or specific semantic units in the text are automatically identified as preliminary candidate features. Clustering threshold parameters are automatically optimized to ensure accurate identification of candidate features.
[0013] Expert knowledge rule base construction combines expert knowledge in the fields of network security, threat intelligence, and log analysis to build and continuously optimize the domain knowledge rule base and clarify the feature attribute classification rules; The candidate feature classification and filtering mechanism automatically classifies candidate features after matching the unsupervised clustering results with the expert knowledge rule base. Successfully matched candidate features are automatically classified, while unsuccessfully matched candidate features are automatically placed in the "pending confirmation" feature base for subsequent manual or semi-supervised confirmation. Feature filtering and quality assessment uses a feature importance scoring mechanism to automatically evaluate the quality of candidate features. Redundant and low-quality candidate features are automatically removed to ensure the accuracy and effectiveness of the extraction results. Storage and management of feature extraction results: Clearly classified and quality-compliant features are uniformly stored in a feature database. The database clearly records the feature source, extraction time, feature type, and associated data for efficient use in subsequent data analysis and event tracing.
[0014] Preferably, the specific process of data cleaning and denoising is: Step W1, missing value processing: eliminate the missing records that cannot be repaired; for the missing data that can be repaired, fill it in through statistical filling method; Step W2, invalid data removal: remove data items that are obviously abnormal, have no analytical value, or have incomplete data records, such as abnormal characters in the log and non-UTF-8 encoded data; Step W3: Data deduplication: Hash the collected data to reduce data redundancy and ease subsequent processing pressure.
[0015] Preferably, UmVAE model training and optimization includes the following steps: Step Q1: training data preparation and processing; Training data: standardized text sequence after binary normalization; Dataset size: no less than 50,000 text data for training and verifying model performance; Step Q2: loss function design and optimization; The loss function is defined as a combination of reconstruction error and KL divergence: ; Reconstruction error: The reconstruction error (cross entropy loss or negative log-likelihood loss) is used to evaluate the model's text reconstruction ability and ensure the semantic consistency between the generated text and the input text; KL divergence: The KL divergence constraint ensures that the distribution of the latent semantic space conforms to the Gaussian distribution, ensuring the generalization ability of the model and ensuring that the semantic space obeys the standard Gaussian distribution; The Adam optimizer is used during training, and the initial learning rate is set to ; Step Q3: Model training process control and evaluation; The model is trained for 30-50 epochs, using early stopping to prevent overfitting. BLEU scores are used for automatic evaluation to ensure consistency between the simplified text and the manually generated standard text. Manual evaluation and confirmation of randomly selected samples ensure accurate semantic expression. When the BLEU score converges stably and the manual evaluation reaches a satisfactory standard, the training is stopped and the optimal model parameters are saved.
[0016] The present invention provides a key attribute extraction method and system based on multimodal normalization. It has the following beneficial effects: This key attribute extraction method and system based on multimodal normalization innovatively integrates unsupervised learning and expert knowledge base, and designs four major modules: data collector, polymorphic element normalizer, complex text simplifier (UmVAE model) and feature extractor. It realizes the automatic normalization of multi-source heterogeneous data, precise semantic simplification and efficient extraction of key attributes, and solves the problems of poor generalization of existing technologies, heavy reliance on labeled data, and low accuracy. It greatly improves data processing efficiency and feature extraction accuracy, and provides strong technical support for correlation analysis, tracing and early warning of data security incidents. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flow chart of the data collector of the present invention; Figure 2 This is a flow chart of the polymorphic element normalizer of the present invention; Figure 3 This is the flow chart of the UmVAE algorithm of the present invention; Figure 4 This is the flow chart of the feature extractor of the present invention. DETAILED DESCRIPTION
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0019] See also Figure 1-4 The present invention provides a technical solution: a key attribute extraction method based on multimodal normalization, which specifically includes the following steps: Step S1: Implementation environment and data preparation; Build an implementation environment within the network security enterprise intranet; Server environment: Two Tesla A100 GPU servers running Python, PyTorch, and the Transformers deep learning framework. Data storage: MongoDB, Elasticsearch, and MinIO are used to build data warehouses and feature databases; Data sources: Threat intelligence platform real-time API interface, network intrusion detection system (IDS) data interface, and terminal log real-time collection interface.
[0020] Step S2: data collection and preliminary processing; Collect unstructured text data from the threat intelligence platform API in real time, and collect network traffic records and terminal system logs from IDS devices; Remove missing records and redundant data, automatically mark data types, and uniformly store data records in JSON format to form a standardized data foundation.
[0021] Step S3, normalization processing of polymorphic elements; Based on the initial data records, a polymorphic element normalizer is deployed to initially establish an expert knowledge base and an automatically updated terminology standardization library to convert text data into unified data.
[0022] Step S4: Simplify the semantics of complex text; Constructing the UmVAE model includes the following steps: Step A1: input normalized data after normalization of polymorphic elements; Step A2: Using unsupervised training, reconstruction error and KL divergence optimization model parameter reparameterization technique to obtain the latent space vector; Step A3: Implement model inference: Input the above normalized data and simplify it through UmVAE to obtain simplified data with precise semantics and clear structure.
[0023] Step S5: feature extraction; Based on the simplified sentences, a feature extractor is deployed to embed the Transformer model features into the semantic vector of the generated data; the DBSCAN unsupervised clustering method is used to automatically identify candidate features, and automatic classification is carried out in combination with the expert knowledge rule base.
[0024] Implementation effect verification: After the implementation was completed, the following results were obtained through actual platform operation verification: the amount of threat intelligence, network traffic and terminal log data processed daily increased from hundreds of thousands to more than one million, the accuracy of text semantic normalization and simplification reached more than 70%, and the accuracy of automatic extraction of key attributes exceeded 62%; the implementation process did not require additional manual data labeling, which significantly reduced labor costs and model iteration cycles, and significantly improved intelligence analysis quality and response efficiency.
[0025] A key attribute extraction system based on multimodal normalization includes a data acquisition unit, a data processing unit and a feature extraction unit; Among them, the data acquisition unit is composed of a real-time data acquisition unit and a data preprocessing unit; The data collection unit collects factual data from the security intelligence platform and threat intelligence library, network security monitoring equipment, and terminal device and server system logs, and cleans and denoises the data. The following is the specific process of data cleaning and denoising: Step W1, missing value processing: eliminate the missing records that cannot be repaired; for the missing data that can be repaired, fill it in through statistical filling method; Step W2, invalid data removal: remove data items that are obviously abnormal, have no analytical value, or have incomplete data records, such as abnormal characters in the log and non-UTF-8 encoded data; Step W3: Data deduplication: Hash the collected data to reduce data redundancy and ease subsequent processing pressure.
[0026] The data preprocessing unit includes preliminary processing of data format normalization, preliminary marking of data labels, and data indexing and storage; Among them, the initial data format normalization process stores heterogeneous data in a unified standard data format, which is conducive to subsequent unified processing. The core information of the original data is extracted and the unified field name and data structure are determined; Preliminary data labeling Preliminary data labels are automatically generated during the preprocessing stage; Data indexing and storage: Use an efficient data indexing engine to index processed data for fast querying. Data storage strategy: It is recommended to store structured data in a time series database and unstructured data in object storage. Clarify storage paths and naming conventions to ensure data orderliness and efficient retrieval.
[0027] The data processing unit includes a pattern matching and replacement module, an automatic construction and iteration knowledge mapping base module, an intelligent abbreviation and non-standard terminology recognition module, and a complex text simplification processing unit; Among them, the pattern matching and replacement module builds a regular expression library to identify non-standard characters; The automatic construction and iteration of the knowledge mapping library module identifies newly emerging non-standard terms and abbreviations from continuously collected data, automatically matches and tags them, and automatically discovers new terms and abbreviations using unsupervised methods. When the frequency of new terms exceeds a threshold, it automatically updates the knowledge base and automatically prompts domain experts for verification and update. After confirmation by domain experts, the knowledge base is automatically updated. The non-standard terminology recognition module uses the BERT pre-trained language model combined with conditional random fields (CRF) to automatically identify abbreviations and non-standard terms. The pre-trained language model is used to encode contextual features of text input; sequence labeling is performed through the CRF model to identify non-standard terms and abbreviations in the text.
[0028] The complex text simplification processing unit is used to build the UmVAE model and train and optimize the constructed UmVAE model; Among them, the UmVAE model consists of three sub-modules: multimodal variational encoder, semantic latent space, and variational decoder; The multimodal variational encoder inputs the binary normalized standard text sequence and uses the Transformer pre-trained language model to convert the input text into a continuous dense vector. The semantic latent space is essentially a low-dimensional dense semantic vector space that can continuously and smoothly represent the semantics of text. During model training, the KL divergence regularization term is used to ensure the structural stability and generalization of the latent semantic space, enabling efficient mapping and conversion from complex text to clear semantics. The variational decoder uses Transformer as the decoder network architecture, takes the vector in the latent space as input, and maps it back to semantically clear and structurally standardized data text.
[0029] UmVAE model training and optimization includes the following steps: Step Q1: Training data preparation and processing Training data: standardized text sequence after binary normalization; Dataset size: no less than 50,000 text data for training and verifying model performance; Step Q2: Loss function design and optimization The loss function is defined as a combination of reconstruction error and KL divergence: ; Reconstruction error: The reconstruction error (cross entropy loss or negative log-likelihood loss) is used to evaluate the model's text reconstruction ability and ensure the semantic consistency between the generated text and the input text; KL divergence: The KL divergence constraint ensures that the distribution of the latent semantic space conforms to the Gaussian distribution, ensures the generalization ability of the model, and ensures that the semantic space obeys the standard Gaussian distribution.
[0030] The Adam optimizer is used during training, and the initial learning rate is set to
[0031] Step Q3: Model training process control and evaluation The model was trained for 30-50 epochs, using early stopping to prevent overfitting. The automatic evaluation metric used the BLEU score to ensure consistency between the simplified text and the manually standardized text. Manual evaluation was performed using randomly selected samples to ensure accurate semantic expression. When the BLEU score converged steadily and the manual evaluation met satisfactory standards, training was stopped and the optimal model parameters were saved.
[0032] The feature extraction unit includes feature embedding and candidate feature identification, expert knowledge rule base construction, candidate feature classification and filtering mechanism; Among them, feature embedding and candidate feature identification are specifically as follows: The input text comes from the standardized text output of the complex text simplifier. A pre-trained Transformer model is used to convert text sentences into contextual semantic feature embedding vectors. An unsupervised semantic clustering algorithm is then used to perform cluster analysis on the embedded features. Frequently occurring phrases, short sentences, or specific semantic units in the text are automatically identified as preliminary candidate features. Clustering threshold parameters are automatically optimized to ensure accurate identification of candidate features.
[0033] The expert knowledge rule base is constructed by combining expert knowledge in the fields of network security, threat intelligence, and log analysis to build and continuously optimize the domain knowledge rule base and clarify the feature attribute classification rules.
[0034] The candidate feature classification and filtering mechanism automatically categorizes candidate features after matching unsupervised clustering results with the expert knowledge rule base. Successfully matched candidate features are automatically classified, while unmatched candidate features are automatically placed in the "pending confirmation" feature library for subsequent manual or semi-supervised confirmation. Feature filtering and quality assessment uses a feature importance scoring mechanism to automatically evaluate the quality of candidate features. Redundant and low-quality candidate features are automatically removed to ensure the accuracy and effectiveness of the extraction results.
[0035] Storage and management of feature extraction results: Clearly classified and quality-compliant features are uniformly stored in a feature database. The database clearly records the feature source, extraction time, feature type, and associated data for efficient use in subsequent data analysis and event tracing.
[0036] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0037] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A key attribute extraction method based on multimodal normalization, The method comprises the following steps: Step S1: Implementation environment and data preparation; Build an implementation environment within the network security enterprise intranet; Data storage: MongoDB, Elasticsearch, and MinIO are used to build data warehouses and feature databases; Data sources: Threat intelligence platform real-time API interface, network intrusion detection system data interface, terminal log real-time collection interface; Step S2: data collection and preliminary processing; Collect unstructured text data from the threat intelligence platform API in real time, and collect network traffic records and terminal system logs from IDS devices; Remove missing records and redundant data, automatically mark data types, and uniformly store data records in JSON format to form a standardized data foundation; Step S3, normalization processing of polymorphic elements; Based on the initial data records, deploy a polymorphic element normalizer to initially establish an expert knowledge base and an automatically updated terminology standardization library to convert text data into unified data; Step S4: Simplify the semantics of complex text and construct a UmVAE model; Step S5: feature extraction; Based on the simplified sentences, a feature extractor is deployed to embed the Transformer model features into the semantic vector of the generated data; the DBSCAN unsupervised clustering method is used to automatically identify candidate features, and automatic classification is carried out in combination with the expert knowledge rule base.
2. The method for extracting key attributes based on multimodal normalization according to claim 1, characterized in that: The step S4 constructs the UmVAE model, specifically comprising the following steps: Step A1: input normalized data after normalization of polymorphic elements; Step A2: Using unsupervised training, reconstruction error and KL divergence optimization model parameter reparameterization technique to obtain the latent space vector; Step A3: Implement model inference: Input the above normalized data and simplify it through UmVAE to obtain simplified data with precise semantics and clear structure.
3. The key attribute extraction system based on multimodal normalization according to claim 1, characterized in that: The data acquisition unit, data processing unit and feature extraction unit; Among them, the data acquisition unit is composed of a real-time data acquisition unit and a data preprocessing unit; The data collection unit collects factual data from the security intelligence platform and threat intelligence database, network security monitoring equipment, and terminal equipment and server system logs, and cleans and removes noise from the data; The data preprocessing unit includes preliminary processing of data format normalization, preliminary marking of data labels, and data indexing and storage; The data format normalization is a preliminary process to store heterogeneous data in a unified standard data format; Preliminary data labeling Preliminary data labels are generated in the preprocessing stage; Data indexing and storage uses an efficient data indexing engine to index the processed data; The data processing unit includes a pattern matching and replacement module, an automatic construction and iteration knowledge mapping base module, an intelligent abbreviation and non-standard terminology recognition module, and a complex text simplification processing unit; Among them, the pattern matching and replacement module builds a regular expression library to identify non-standard characters; The automatic construction and iteration of the knowledge mapping library module identifies newly emerging non-standard terms and abbreviations from continuously collected data, automatically matches and tags them, and automatically discovers new terms and abbreviations using unsupervised methods; The non-standard term recognition module uses the BERT pre-trained language model and combines it with conditional randomization to automatically recognize abbreviations and non-standard terms. The complex text simplification processing unit is used to build the UmVAE model and train and optimize the constructed UmVAE model; The feature extraction unit includes feature embedding and candidate feature identification, expert knowledge rule base construction, candidate feature classification and filtering mechanism; The feature embedding and candidate feature identification process involves the following steps: the input text comes from the normalized text output by the complex text simplifier, the Transformer pre-trained model is used to convert the text sentences into contextual semantic feature embedding vectors, and an unsupervised semantic clustering algorithm is used to perform cluster analysis on the embedded features; Expert knowledge rule base construction combines expert knowledge in the fields of network security, threat intelligence, and log analysis to build a domain knowledge rule base; The candidate feature classification and filtering mechanism automatically classifies the candidate features after matching the unsupervised clustering results with the expert knowledge rule base, and automatically classifies the successfully matched candidate features; The storage and management of feature extraction results will uniformly store features with clear classification and qualified quality in the feature database.
4. The key attribute extraction system based on multimodal normalization according to claim 3 is characterized by: The UmVAE model consists of three submodules: multimodal variational encoder, semantic latent space, and variational decoder; The multimodal variational encoder inputs the binary normalized standard text sequence and uses the Transformer pre-trained language model to convert the input text into a continuous dense vector. The semantic latent space is specifically a low-dimensional dense semantic vector space that continuously and smoothly represents the semantics of the text; The variational decoder uses Transformer as the decoder network architecture, takes the vector in the latent space as input, and maps it back to semantically clear and structurally standardized data text.
5. The key attribute extraction system based on multimodal normalization according to claim 3 is characterized by: The specific process of data cleaning and denoising is as follows: Step W1, missing value processing: eliminate the missing records that cannot be repaired; for the missing data that can be repaired, fill it in through statistical filling method; Step W2, invalid data removal: remove data items that are obviously abnormal, have no analytical value, or have incomplete data records, such as abnormal characters in the log and non-UTF-8 encoded data; Step W3: Data deduplication: Hash the collected data to reduce data redundancy and ease subsequent processing pressure.
6. The key attribute extraction system based on multimodal normalization according to claim 3, characterized in that: The UmVAE model training and optimization includes the following steps: Step Q1: training data preparation and processing; Training data: standardized text sequence after binary normalization; Dataset size: no less than 50,000 text data for training and verifying model performance; Step Q2: loss function design and optimization; The loss function is defined as a combination of reconstruction error and KL divergence: ; Reconstruction error: The reconstruction error is used to evaluate the model's text reconstruction ability and ensure the semantic consistency between the generated text and the input text; KL divergence: The KL divergence constraint ensures that the distribution of the latent semantic space conforms to the Gaussian distribution, ensuring the generalization ability of the model and ensuring that the semantic space obeys the standard Gaussian distribution; The Adam optimizer is used during training, and the initial learning rate is set to ; Step Q3: Model training process control and evaluation; The model is trained for 30-50 rounds, using early stopping to prevent overfitting, and the automatic evaluation metric uses BLEU score; When the BLEU score converges stably and the manual evaluation reaches a satisfactory standard, the training is stopped and the optimal model parameters are saved.
Citation Information
Patent Citations
Multi-modal knowledge graph construction method
CN112200317A
Automatic discovery and evolution method and system for terminology and research field, terminal and medium
CN117195874A
Method for mining and analyzing abnormal security behaviors of industrial internet
CN119030799A
Knowledge graph construction method and system based on large model technology
CN119494390A
Network security risk identification and management and control system based on AI
CN119496647A
Cited By
Data extraction method based on data characteristics of signal system, medium and electronic equipment
CN121579701A