Cross-modal image and text corpus association analysis system

Through deep learning models and cross-modal correlation learning methods, the adaptability and semantic consistency of the correlation analysis of cross-modal images and text corpus in the power industry in the existing technology is solved, and higher analysis accuracy and fault warning accuracy are achieved.

CN120337129APending Publication Date: 2025-07-18CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510389696.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art cross-modal image and text corpus correlation analysis methods are poorly adaptable in the power industry, making it difficult to effectively capture complex features and semantic consistency, resulting in insufficient accuracy of fault warning.

Method used

Deep learning model is used for feature extraction, image and text features are extracted through convolutional neural network (CNN) and natural language processing (NLP), and feature alignment, weighted summing and attention mechanisms are combined with cross-modal association learning modules to enhance semantic consistency, and unsupervised learning is used to optimize model parameters.

Benefits of technology

It improves the accuracy and reliability of image-text correlation analysis, enhances the ability to capture complex features, and improves the accuracy of power equipment fault diagnosis and early warning.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention provides a cross-modal image and text corpus association analysis system, and relates to the field of image and text analysis. Comprising a data collection and preprocessing module, a feature extraction module, a cross-modal association learning module and a model training and optimization module. The data collection and preprocessing module collects image and text data from multiple sources and preprocesses the image and text data; the feature extraction module extracts image and text features by using CNN and NLP models; the cross-modal association learning module enhances the semantic consistency of image and text features through feature alignment, weighted summation, an attention mechanism and a cross-modal interaction unit; the model training and optimizing module adopts an unsupervised learning method to train a model and uses an optimization algorithm to adjust parameters; according to the system, through an innovative cross-modal association learning mechanism and an advanced deep learning model, the accuracy and reliability of cross-modal image and text corpus association analysis are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image and text analysis, and specifically to a cross-modal image and text corpus correlation analysis system. Background Art

[0002] With the rapid development of artificial intelligence and big data technologies, cross-modal data analysis plays an important role in fields such as smart grid monitoring, power equipment fault diagnosis, and transmission line inspection analysis, becoming a hot research field; it can effectively integrate different types of data resources, mine deeper information, which is of great significance for improving system performance and decision-making accuracy, and is expected to be widely applied in more fields in the future.

[0003] Currently, common technical solutions for cross-modal image and text corpus correlation analysis mainly adopt traditional feature extraction methods. For example, image processing is performed using manually designed feature extraction operators, and text feature representation is carried out using a simple bag-of-words model, and then correlation analysis is achieved through some basic similarity measurement methods; the working principle of such solutions is relatively simple, mainly focusing on basic feature extraction and matching, characterized by low implementation difficulty and low requirements for computing resources.

[0004] The deficiencies of the prior art are that traditional feature extraction methods have poor adaptability to power industry data, are difficult to effectively capture the deep correlations of complex features such as equipment defects and thermal anomalies, resulting in insufficient accuracy of fault warnings and limited accuracy of correlation analysis; moreover, basic similarity measurement methods lack in-depth exploration of semantic consistency and cannot fully consider the complex correlation relationships between images and texts; in view of these problems, this paper proposes a new cross-modal image and text corpus correlation analysis system, which can more accurately capture the deep features and semantic consistency of images and texts by introducing advanced deep learning models for feature extraction and adopting an innovative cross-modal correlation learning mechanism, thereby effectively improving the accuracy and reliability of correlation analysis. Summary of the Invention

[0005] (1) Technical Problems to be Solved

[0006] Aiming at the deficiencies of the prior art, the present invention provides a cross-modal image and text corpus correlation analysis system to solve the problems that traditional feature extraction methods have poor data adaptability and are difficult to effectively capture the deep features of complex data, resulting in limited accuracy of correlation analysis, and that basic similarity measurement methods lack in-depth exploration of semantic consistency and cannot fully consider the complex correlation relationships between images and texts as proposed in the above background art.

[0007] (2) Technical Solutions

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: A cross-modal image and text corpus correlation analysis system, including a data collection and preprocessing module, a feature extraction module, a cross-modal correlation learning module, and a model training and optimization module; wherein:

[0009] The data collection and preprocessing module is used to collect image data and text data from multiple data sources, and perform size adjustment, normalization processing, and data augmentation on the images, and perform word segmentation, stop word removal, and lemmatization processing on the text data;

[0010] The feature extraction module extracts image features through a convolutional neural network (CNN) model and extracts text features through a natural language processing (NLP) model;

[0011] The cross-modal correlation learning module fuses image and text features through feature alignment, weighted summation, and attention mechanism, and further enhances the semantic consistency of image and text features through a cross-modal interaction unit;

[0012] The model training and optimization module trains the model through unsupervised learning and uses an optimization algorithm to adjust the model parameters to optimize the semantic correlation between image and text features.

[0013] Preferably, the data collection and preprocessing module collects image data and text data from multiple data sources. The sources of image data include the Internet, power industry data platforms, and professional image databases, where the types of image data cover transmission line inspection pictures, product pictures, art works, and satellite images; the collected text data includes equipment maintenance logs, fault diagnosis reports, operation manuals, inspection records, power dispatching instructions, and technical specification documents related to the images, and obtains annotation data between images and texts according to requirements; in the image preprocessing step, the image size is first adjusted to a unified size, specifically 224×224 pixels; then the pixel values of the image are normalized, and the normalization range can be selected as [0,1] or [-1,1]; subsequently, data augmentation techniques such as rotation, cropping, and flipping operations are used to increase data diversity; in addition, the image is denoised through Gaussian filtering to improve data quality; text preprocessing includes using a word segmentation tool to split the text into sub-word units, removing stop words from the text, and performing lemmatization processing; finally, pre-training embedding is performed using the BERT model to convert the text into a numerical vector representation; the word segmentation tool uses the Word2Vec model for sub-word segmentation, where the stop word list can be customized according to specific application scenarios; lemmatization is implemented using the NLTK natural language processing library; the BERT model pre-training embedding maps the text sequence to a 768-dimensional vector space by loading the pre-trained BERT weights.

[0014] Preferably, in the feature extraction module, image feature extraction is performed through a pre-trained convolutional neural network (CNN) model. The global features of the image are extracted through the global average pooling layer of the CNN. The global average pooling layer performs global average pooling on the feature map output by the convolutional layer to obtain a feature vector of a fixed length. The local features of the image are extracted through the convolutional layer of the CNN. The convolutional layer extracts the features of local regions on the image by means of a sliding window to capture the local information in the image. The text feature extraction is performed through a pre-trained natural language processing (NLP) model, namely, the BERT model. The BERT model generates the semantic embedding of the entire sentence through the CLS token, and the dimension representing the sentence embedding is 768. The word embeddings in the text are extracted through the Word2Vec pre-trained model to convert words into vector representations of a fixed dimension. The CNN model can adopt the Res Net pre-trained model for feature extraction. The Word2Vec model learns the distributed representation of words by training a large amount of text data and maps each word into a vector space of a fixed dimension.

[0015] Preferably, the cross-modal association learning module uses feature alignment to map image and text features into a feature space of the same dimension through a non-linear transformation method. The non-linear transformation is performed through the Transformer model. The Transformer model uses the self-attention mechanism to perform weighted summation on the input features to achieve non-linear transformation of the features. The multi-modal fusion is performed through weighted summation and the attention mechanism. The weighted summation adjusts the weights according to the relative importance of the image and text features, and the weights are obtained through training optimization. The attention mechanism dynamically adjusts the weights of the image and text features to ensure that important features are highlighted. The cross-modal interaction unit enables mutual influence and fusion between image and text features through the use of the Cross-Attention method, enhancing the semantic consistency of the image and text features. The Cross-Attention method realizes the interaction and fusion of features by calculating the attention weights between image features and text features. The feature selection is performed by calculating the mutual information content of the image and text features to ensure that the final joint feature space has high expressive power and low redundancy. The mutual information content is obtained by calculating the logarithmic ratio of the joint probability distribution and the marginal probability distribution between features and performing integration.

[0016] Preferably, the model training and optimization module uses an unsupervised learning method for training; the unsupervised learning task is contrastive learning, and the model is trained through a contrastive loss function to maximize the matching degree between images and texts; then the Adam optimization algorithm is used for parameter update, and the Adam optimizer adaptively adjusts the learning rate to make the model converge faster; during the training process, a regularization method is used to prevent overfitting, and an early stopping technique is used to monitor the loss of the validation set to avoid performance degradation caused by overtraining; the Adam optimizer calculates the first-order moment estimate and second-order moment estimate of the gradient; and adaptively adjusts the learning rate to accelerate the convergence of the model; through the contrastive learning method, the unsupervised learning enables the model to learn the feature representations of images and texts without labeled data, thereby improving the generalization ability and adaptability of the model.

[0017] Preferably, the data collection and preprocessing module further uses a cross-validation method to partition and optimize the weights of the dataset during the collection of image and text data to ensure the accuracy of the matching relationship between images and texts; among them, the image data is denoised to improve the data quality, and the text data ensures the semantic consistency of texts from different sources through unified word segmentation and lemmatization; the cross-validation method divides the dataset into k subsets, and each time k - 1 subsets are used for training, and the remaining 1 subset is used for validation. The average performance index of the model is calculated through multiple iterations to optimize the generalization ability of the model; the denoising process smooths the image through Gaussian filtering, removes the noise points in the image, and improves the clarity of the image; the unified word segmentation and lemmatization standardize the text through the NLTK natural language processing library; to ensure the semantic consistency of texts from different sources.

[0018] Preferably, the CNN model in the feature extraction module uses a pre-trained Res Net model for image feature extraction; the Res Net model solves the problem of gradient disappearance in the training of deep networks through residual connections and can effectively extract the deep features of images; the BERT model uses a pre-trained BERT-base model for text feature extraction, and the BERT-base model contains 12 layers of Transformer encoders and 110M parameters and can capture long-distance dependencies and semantic information in texts; the Word2Vec model uses a pre-trained Google News Word2Vec model for word embedding, and the Word2Vec model learns the distributed representation of words by training data of 100 billion words, enabling it to effectively capture the semantic relationships between words.

[0019] Preferably, the feature alignment in the cross-modal association learning module maps image and text features to a feature space of the same dimension through a non-linear transformation method; the non-linear transformation is implemented by the self-attention mechanism in the Transformer model. The self-attention mechanism performs weighted summation of features by calculating the correlation between features, thereby realizing the non-linear transformation of features; the weighted summation adjusts the weights according to the relative importance of image and text features, and the weights are obtained by calculating the similarity matrix of image and text features and normalizing it using the softmax function; the attention mechanism dynamically adjusts the weights of image and text features to ensure that important features are highlighted, and the attention weights are obtained by calculating the dot product between features and normalizing it using the softmax function.

[0020] Preferably, the dynamic learning rate scheduler in the model training and optimization module gradually reduces the learning rate with the progress of training through cosine annealing, piecewise constant, and exponential decay methods; the cosine annealing method simulates the periodic change of the cosine function to gradually reduce the learning rate during training; the piecewise constant method divides the training process into multiple stages, where the learning rate remains constant within each stage, but gradually decreases between stages; the exponential decay method gradually reduces the learning rate during training through an exponential function; the validation set evaluation evaluates the performance of the model by calculating performance metrics of the model on the validation set, such as accuracy, recall, and F1 value, etc., and adjusts the model parameters and training strategy according to the evaluation results.

[0021] Preferably, the cross-modal association modeling based on causal reasoning in the cross-modal association learning module constructs a causal graph model between images and texts through a causal discovery algorithm; the causal reasoning model uses the do-calculus method for reasoning; the reasoning process simulates the causal impact of image or text changes on the other modality, thereby improving the interpretability of the association between images and texts; the causal graph is embedded in the parameter space of the neural network through the variational causal encoder VCE to achieve end-to-end training, enabling the model to learn causal relationships and feature representations simultaneously during training, thereby optimizing the robustness and accuracy of cross-modal analysis; the causal discovery algorithm adopts the PC algorithm; the causal graph is constructed through conditional independence testing; the do-calculus method intervenes in causal variables through the do operator to calculate the causal effect.

[0022] (III) Beneficial Effects

[0023] The present invention provides a cross-modal image and text corpus association analysis system. It has the following beneficial effects:

[0024] 1. Through the methods of feature alignment and multimodal fusion, the present invention maps image and text features into the same feature space, and performs fusion through weighted summation and attention mechanism, enhancing the semantic consistency of image and text features. Specifically, feature alignment adopts non-linear transformation methods, such as the self-attention mechanism in the Transformer model, which can effectively capture the correlation between features and achieve complex non-linear transformation. Weighted summation adjusts the weights according to the relative importance of image and text features, and the weights are obtained through training optimization to ensure that important features are highlighted. The attention mechanism further enhances the semantic consistency of features by dynamically adjusting feature weights.

[0025] 2. The present invention uses unsupervised learning methods for model training. Through methods such as contrastive learning, autoencoders, cluster adversarial training, and joint training, the model can learn the feature representations of images and texts without labeled data. Unsupervised learning methods can not only make full use of a large amount of unlabeled data, but also improve the generalization ability and adaptability of the model. And through strategies such as data augmentation, cross-validation, adversarial training, data standardization and normalization, adaptive learning strategies, and multimodal information fusion, it is further ensured that the model can perform well on multiple data sets. Specific Embodiments

[0026] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] Embodiment 1:

[0028] The embodiment of the present invention provides a cross-modal image and text corpus correlation analysis system. The specific implementation process is as follows: First, in the data collection and preprocessing module, image data and text data are collected from multiple data sources. The sources of image data include the Internet power industry data platform and professional image databases. The types of image data cover transmission line inspection pictures, commodity pictures, art works, and satellite images. The text data includes titles, descriptions, tags, and operation and maintenance records related to the images, and annotation data between the images and texts is obtained according to requirements. Specifically, data interfaces are used to obtain image data from public databases in the power industry, such as the State Grid equipment picture library and the internal systems of power equipment manufacturers, and transmission line thermal imaging maps and substation monitoring images are downloaded through the power grid inspection system. For text data, equipment maintenance logs and fault reports are obtained through the internal document system of the power company, and technical specification documents and operation manuals are extracted from the power industry standard library.

[0029] In the image preprocessing step, the image size is adjusted to 224×224 pixels, and the Pillow (PIL) library is used for size adjustment. Then, the pixel values of the image are normalized within the range [0, 1], and the numpy library is used for pixel value normalization. Subsequently, data augmentation techniques such as rotation, cropping, and flipping operations are used to increase data diversity, and the OpenCV library is used to implement these data augmentation operations.

[0030] In addition, the image is denoised by Gaussian filtering to improve data quality, and the gaussian_filter function in the SciPy library is used for denoising. In text preprocessing, a tokenizer is used to split the text into sub-word units, stop words in the text are removed, and lemmatization is performed. Then, the BERT model is used for pre-training embedding to convert the text into a numerical vector representation. Specifically, the NLTK library is used for tokenization and lemmatization, and the Hugging Face Transformers library is used to load the pre-trained BERT model for text embedding.

[0031] In the feature extraction module, image features are extracted by a pre-trained convolutional neural network (CNN) model, and text features are extracted by a pre-trained natural language processing (NLP) model, namely the BERT model. For image feature extraction, the pre-trained ResNet model is used to extract the global features of the image through the global average pooling layer and the local features of the image through the convolutional layer. For text feature extraction, the BERT model is used to generate the semantic embedding of the entire sentence through the CLS token, and the dimension of the sentence embedding is 768. The word embeddings in the text are extracted by the Word2Vec pre-trained model to convert words into vector representations of a fixed dimension.

[0032] In the cross-modal association learning module, feature alignment is used to map image and text features to a feature space of the same dimension through a non-linear transformation method. The non-linear transformation is performed by the Transformer model. Multi-modal fusion is carried out through weighted summation and attention mechanism. The weighted summation adjusts the weights according to the relative importance of image and text features, and the weights are optimized through training. The attention mechanism dynamically adjusts the weights of image and text features to ensure that important features are highlighted. The cross-modal interaction unit uses the Cross-Attention method to enable mutual influence and fusion between image and text features, enhancing the semantic consistency of image and text features.

[0033] In the model training and optimization module, an unsupervised learning method is adopted for training. The unsupervised learning task is contrastive learning. The model is trained through a contrastive loss function to maximize the semantic matching degree between the power equipment images and the corresponding operation and maintenance texts, supporting equipment status prediction and fault cause analysis. The Adam optimization algorithm is used for parameter update. The Adam optimizer enables the model to converge faster by adaptively adjusting the learning rate. During the training process, regularization methods are used to prevent overfitting, and early stopping techniques are used to monitor the validation set loss to avoid performance degradation caused by overtraining and ensure that the model has strong generalization ability. Specifically, the collected image and text data are divided into a training set, a validation set, and a test set according to a ratio of 7:2:1. The model is trained on the training set, the model is evaluated and parameter adjusted using the validation set, and the final model performance test is conducted on the test set. Through the above steps, the implementation of the cross-modal image and text corpus association analysis system is completed.

[0034] Embodiment 2:

[0035] This embodiment further optimizes the data collection and preprocessing module and the model training and optimization module on the basis of Embodiment 1. In the data collection and preprocessing module, a cross-validation method for image and text data is added to divide and optimize the weights of the data set to ensure the accuracy of the matching relationship between images and texts. Specifically, the data set is divided into 5 subsets. Each time, 4 subsets are used for training, and the remaining 1 subset is used for validation. The average performance index of the model is calculated through multiple iterations to optimize the generalization ability of the model. In image preprocessing, a denoising process step is added. The image is smoothed through Gaussian filtering to remove the noise points in the image to improve the clarity of the image. In text preprocessing, steps of unified word segmentation and lemmatization are added. The text is standardized using the NLTK natural language processing library to ensure the semantic consistency of texts from different sources.

[0036] A dynamic learning rate scheduler is added to the model training and optimization module. The learning rate is gradually decreased with the training process through the cosine annealing piecewise constant and exponential decay methods. Specifically, the learning rate is relatively high at the beginning of training and gradually decreases as training progresses, so as to improve the convergence speed and stability of the model. In the validation set evaluation, detailed evaluation metrics for model performance, such as accuracy, recall, F1 value, etc., are added, and the model parameters and training strategies are adjusted according to the evaluation results. In addition, in the cross-modal association learning module, cross-modal association modeling based on causal reasoning is added. A causal graph model between images and texts is constructed through a causal discovery algorithm. The causal reasoning model uses the do-calculus method for reasoning. The reasoning process simulates the changes in power equipment images, such as the causal impact of insulator damage on text descriptions, or the feedback effect of text instructions, i.e., dispatching commands, on device operation images, enhancing the model's ability to explain the causal relationships in the power system, thereby improving the interpretability of the association between images and texts. The causal graph is embedded in the parameter space of the neural network through the variational causal encoder VCE to achieve end-to-end training, enabling the model to learn causal relationships and feature representations simultaneously during training, thus optimizing the robustness and accuracy of cross-modal analysis. Through the above optimization steps, the performance and accuracy of the cross-modal image and text corpus association analysis system are further improved.

[0037] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A cross-modal image and text corpus correlation analysis system, characterized in that It includes a data collection and preprocessing module, a feature extraction module, a cross-modal correlation learning module, and a model training and optimization module; among which: The data collection and preprocessing module is used to collect image data and text data from multiple data sources, perform size adjustment, normalization, denoising, and data augmentation on the images, perform word segmentation, stop word removal, and lemmatization on the text data, and use the BERT model for pre-training embedding; The feature extraction module extracts image features through a convolutional neural network (CNN) model and extracts text features through a natural language processing (NLP) model; The cross-modal correlation learning module fuses image and text features through feature alignment, weighted summation, and attention mechanism, and further enhances the semantic consistency of image and text features through a cross-modal interaction unit; The model training and optimization module trains the model through unsupervised learning and uses an optimization algorithm to adjust the model parameters to optimize the semantic association between image and text features.

2. The cross-modal image and text corpus correlation analysis system according to claim 1, wherein: The data collection and preprocessing module collects image data from multiple data sources. The sources of the image data include public databases in the power industry, galleries of power equipment manufacturers, and inspection systems of power grid companies. The types of image data include substation equipment images, transmission line inspection pictures, thermal imaging maps of power equipment, power grid topology diagrams, and distribution facility monitoring images; among which, the text data related to the images includes titles, descriptions, tags, and operation and maintenance records, and annotation data between images and texts is obtained according to requirements; the image preprocessing steps include adjusting the image size to a unified size, specifically 224×224 pixels; normalizing the image pixel values, and the normalization range is [0,1] or [-1,1]; the denoising process removes noise in the image by using Gaussian filtering; and data augmentation techniques are used, including rotation, cropping, and flipping operations; the text data is segmented into sub-word units by using a word segmentation tool, stop words in the text are removed, lemmatization is performed, and the BERT model is used for pre-training embedding to convert the text into a numerical vector representation.

3. The cross-modal image and text corpus correlation analysis system according to claim 1, characterized in that: In the feature extraction module, image feature extraction is performed through a pre-trained convolutional neural network CNN model. Among them, the global features of the image are extracted through the global average pooling layer of the CNN; the local features of the image are extracted through the convolutional layer of the CNN to capture local region information in the image; text feature extraction is performed through a pre-trained natural language processing NLP model. The BERT model generates a semantic embedding representation of the entire sentence through the CLS token, and the dimension of the sentence embedding is 768; the word embeddings in the text are extracted through the Word2Vec pre-trained model to convert words into vector representations of a fixed dimension.

4. A cross-modal image and text corpus correlation analysis system according to claim 1, characterized in that: The cross-modal correlation learning module uses feature alignment through a non-linear transformation method to map image and text features into a feature space of the same dimension. The non-linear transformation is performed through a Transformer model. Among them, multi-modal fusion is carried out through weighted summation and attention mechanism. The weighted summation adjusts the weights according to the relative importance of image and text features, and the weights are obtained through training optimization. The attention mechanism dynamically adjusts the weights of image and text features to ensure that important features are highlighted. The cross-modal interaction unit uses the Cross-Attention method to achieve semantic alignment between power equipment image features and text features, enhancing the correlation between equipment anomalies and text descriptions. The power equipment image features are thermal imaging maps and equipment status maps, and the text features are fault diagnosis reports and maintenance logs.

5. A cross-modal image and text corpus correlation analysis system according to claim 1, characterized in that: The cross-modal correlation learning module further optimizes the joint representation of image and text features through a feature selection method, and the feature selection is carried out by calculating the mutual information of image and text features.

6. The cross-modal image and text corpus correlation analysis system according to claim 1, characterized in that: In terms of the training method, the model training and optimization module specifically uses an unsupervised learning method for training. The unsupervised learning task is contrastive learning, and the model is trained through a contrastive loss function to maximize the matching degree between images and texts. Then, the Adam optimization algorithm is used to update the parameters. The Adam optimizer adaptively adjusts the learning rate to make the model converge faster. During the training process, a regularization method is used to prevent overfitting, and an early stopping technique is used to monitor the validation set loss to avoid performance degradation caused by over-training.

7. A cross-modal image and text corpus correlation analysis system according to claim 1, characterized in that: The model training and optimization module also includes adjusting the learning rate through a dynamic learning rate scheduler during the training process. The learning rate scheduler gradually reduces the learning rate with the progress of training through methods such as cosine annealing, piecewise constant, and exponential decay. After the training is completed, the validation set is used to evaluate the model performance. According to the evaluation results, the model parameters and training strategies are adjusted to ensure that the model achieves the best generalization ability and can handle different data sources and task scenarios.

8. A cross-modal image and text corpus correlation analysis system according to claim 1, characterized in that: The cross-modal correlation learning module also includes cross-modal correlation modeling based on causal reasoning. A causal graph model between images and texts is constructed through a causal discovery algorithm. The causal reasoning model uses the do-calculus method for reasoning, and the reasoning process simulates the causal impact of image or text changes on the other modality, thereby improving the interpretability of the correlation between images and texts. The causal graph is embedded in the parameter space of the neural network through a variational causal encoder (VCE) to achieve end-to-end training, enabling the model to learn causal relationships and feature representations simultaneously during the training process, thereby optimizing the robustness and accuracy of cross-modal analysis.

9. The cross-modal image and text corpus correlation analysis system according to claim 1, wherein: During the collection of image and text data, the data collection and preprocessing module further uses a cross-validation method to divide and optimize the weights of the dataset to ensure the accuracy of the matching relationship between images and texts. Among them, the image data is processed by denoising to improve the data quality, and the text data is ensured to have semantic consistency for texts from different sources through unified word segmentation and lemmatization.

Citation Information

Cited By

  • Intelligent service execution method based on dialogue mechanism

    CN121009162A

  • Document analysis evaluation method and system based on multi-modal semantic consistency

    CN121052242A

  • Intelligent data annotation method and system based on cross-modal joint learning

    CN121786763A